VLDB 2026 Research / reviewers in the wild / expert
Ziwei Ji 0001
dblp:176/4574-1
· DBLP profile ↗
10ranked-venue papers
2as first author
10since 2021 · last 2025
0000-0002-0206-7861ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HalluLens: LLM Hallucination BenchmarkabstractYejin Bang, Ziwei Ji, Alan Schelten, Anthony Hartshorn, Tara Fowler, Cheng Zhang, Nicola Cancedda, Pascale Fung. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yejin Bang, Ziwei Ji 0001, Alan Schelten, Anthony Hartshorn, Tara Fowler, Nicola Cancedda, Pascale Fung |
ACL (1) | 2 |
| 2025 | Calibrating Verbal Uncertainty as a Linear Feature to Reduce HallucinationsabstractZiwei Ji, Lei Yu, Yeskendir Koishekenov, Yejin Bang, Anthony Hartshorn, Alan Schelten, Cheng Zhang, Pascale Fung, Nicola Cancedda. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Ziwei Ji 0001, Yeskendir Koishekenov, Yejin Bang, Anthony Hartshorn, Alan Schelten, Pascale Fung, Nicola Cancedda |
EMNLP | 1 |
| 2025 | High-Dimension Human Value Representation in Large Language ModelsabstractSamuel Cahyawijaya, Delong Chen, Yejin Bang, Leila Khalatbari, Bryan Wilie, Ziwei Ji, Etsuko Ishii, Pascale Fung. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Samuel Cahyawijaya, Delong Chen, Yejin Bang, Leila Khalatbari, Bryan Wilie, Ziwei Ji 0001, Etsuko Ishii, Pascale Fung |
NAACL (Long Papers) | 6 |
| 2024 | ANAH: Analytical Annotation of Hallucinations in Large Language ModelsabstractReducing the 'hallucination' problem of Large Language Models (LLMs) is crucial for their wide applications.A comprehensive and finegrained measurement of the hallucination is the first key step for the governance of this issue but is under-explored in the community.Thus, we present ANAH, a bilingual dataset that offers ANalytical Annotation of Hallucinations in LLMs within Generative Question Answering.Each answer sentence in our dataset undergoes rigorous annotation, involving the retrieval of a reference fragment, the judgment of the hallucination type, and the correction of hallucinated content.ANAH consists of ∼12k sentence-level annotations for ∼4.3kLLM responses covering over 700 topics, constructed by a human-in-the-loop pipeline.Thanks to the fine granularity of the hallucination annotations, we can quantitatively confirm that the hallucinations of LLMs progressively accumulate in the answer and use ANAH to train and evaluate hallucination annotators.We conduct extensive experiments on studying generative and discriminative annotators and show that, although current open-source LLMs have difficulties in fine-grained hallucination annotation, the generative annotator trained with ANAH can surpass all open-source LLMs and GPT-3.5, obtain performance competitive with GPT-4, and exhibits better generalization ability on unseen questions. 1 Ziwei Ji 0001, Yuzhe Gu, Chengqi Lyu, Dahua Lin, Kai Chen 0026 |
ACL (1) | 1 |
| 2024 | ANAH-v2: Scaling Analytical Hallucination Annotation of Large Language ModelsabstractLarge language models (LLMs) exhibit hallucinations in long-form question-answering tasks across various domains and wide applications. Current hallucination detection and mitigation datasets are limited in domain and size, which struggle to scale due to prohibitive labor costs and insufficient reliability of existing hallucination annotators. To facilitate the scalable oversight of LLM hallucinations, this paper introduces an iterative self-training framework that simultaneously and progressively scales up the annotation dataset and improves the accuracy of the annotator. Based on the Expectation Maximization algorithm, in each iteration, the framework first applies an automatic hallucination annotation pipeline for a scaled dataset and then trains a more accurate annotator on the dataset. This new annotator is adopted in the annotation pipeline for the next iteration. Extensive experimental results demonstrate that the finally obtained hallucination annotator with only 7B parameters surpasses GPT-4 and obtains new state-of-the-art hallucination detection results on HaluEval and HalluQA by zero-shot inference. Such an annotator can not only evaluate the hallucination levels of various LLMs on the large-scale dataset but also help to mitigate the hallucination of LLMs generations, with the Natural Language Inference metric increasing from 25% to 37% on HaluEval. Yuzhe Gu, Ziwei Ji 0001, Chengqi Lyu, Dahua Lin, Kai Chen 0026 |
NeurIPS | 2 |
| 2023 | Plausible May Not Be Faithful: Probing Object Hallucination in Vision-Language Pre-trainingabstractLarge-scale vision-language pre-trained (VLP) models are prone to hallucinate non-existent visual objects when generating text based on visual information.In this paper, we systematically study the object hallucination problem from three aspects.First, we examine recent state-of-the-art VLP models, showing that they still hallucinate frequently, and models achieving better scores on standard metrics (e.g., CIDEr) could be more unfaithful.Second, we investigate how different types of image encoding in VLP influence hallucination, including region-based, grid-based, and patch-based.Surprisingly, we find that patch-based features perform the best and smaller patch resolution yields a non-trivial reduction in object hallucination.Third, we decouple various VLP objectives and demonstrate that token-level imagetext alignment and controlled generation are crucial to reducing hallucination.Based on that, we propose a simple yet effective VLP loss named ObjMLM to further mitigate object hallucination.Results show that it reduces object hallucination by up to 17.4% when tested on two benchmarks (COCO Caption for in-domain and NoCaps for out-of-domain evaluation). Wenliang Dai, Zihan Liu 0001, Ziwei Ji 0001, Dan Su 0003, Pascale Fung |
EACL | 3 |
| 2023 | Contrastive Learning for Inference in DialogueabstractInference, especially those derived from inductive processes, is a crucial component in our conversation to complement the information implicitly or explicitly conveyed by a speaker.While recent large language models show remarkable advances in inference tasks, their performance in inductive reasoning, where not all information is present in the context, is far behind deductive reasoning.In this paper, we analyze the behavior of the models based on the task difficulty defined by the semantic information gap -which distinguishes inductive and deductive reasoning (Johnson-Laird, 1988, 1993).Our analysis reveals that the disparity in information between dialogue contexts and desired inferences poses a significant challenge to the inductive inference process.To mitigate this information gap, we investigate a contrastive learning approach by feeding negative samples.Our experiments suggest negative samples help models understand what is wrong and improve their inference generations. 1 Etsuko Ishii, Yan Xu 0012, Bryan Wilie, Ziwei Ji 0001, Holy Lovenia, Willy Chung, Pascale Fung |
EMNLP | 4 |
| 2023 | Diverse and Faithful Knowledge-Grounded Dialogue Generation via Sequential Posterior InferenceabstractThe capability to generate responses with diversity and faithfulness using factual knowledge is paramount for creating a human-like, trustworthy dialogue system. Common strategies either adopt a two-step paradigm, which optimizes knowledge selection and response generation separately, and may overlook the inherent correlation between these two tasks, or leverage conditional variational method to jointly optimize knowledge selection and response generation by employing an inference network. In this paper, we present an end-to-end learning framework, termed Sequential Posterior Inference (SPI), capable of selecting knowledge and generating dialogues by approximately sampling from the posterior distribution. Unlike other methods, SPI does not require the inference network or assume a simple geometry of the posterior distribution. This straightforward and intuitive inference procedure of SPI directly queries the response generation model, allowing for accurate knowledge selection and generation of faithful responses. In addition to modeling contributions, our experimental results on two common dialogue datasets (Wizard of Wikipedia and Holl-E) demonstrate that SPI outperforms previous strong baselines according to both automatic and human evaluation metrics. Yan Xu 0012, Deqian Kong, Dehong Xu, Ziwei Ji 0001, Bo Pang 0004, Pascale Fung, Ying Nian Wu |
ICML | 4 |
| 2023 | A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and InteractivityabstractYejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, Quyet V. Do, Yan Xu, Pascale Fung. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su 0003, Bryan Wilie, Holy Lovenia, Ziwei Ji 0001, Tiezheng Yu, Willy Chung, Quyet V. Do, Yan Xu 0012, Pascale Fung |
IJCNLP (1) | 8 |
| 2021 | CrossNER: Evaluating Cross-Domain Named Entity RecognitionabstractCross-domain named entity recognition (NER) models are able to cope with the scarcity issue of NER samples in target domains. However, most of the existing NER benchmarks lack domain-specialized entity types or do not focus on a certain domain, leading to a less effective cross-domain evaluation. To address these obstacles, we introduce a cross-domain NER dataset (CrossNER), a fully-labeled collection of NER data spanning over five diverse domains with specialized entity categories for different domains. Additionally, we also provide a domain-related corpus since using it to continue pre-training language models (domain-adaptive pre-training) is effective for the domain adaptation. We then conduct comprehensive experiments to explore the effectiveness of leveraging different levels of the domain corpus and pre-training strategies to do domain-adaptive pre-training for the cross-domain task. Results show that focusing on the fractional corpus containing domain-specialized entities and utilizing a more challenging pre-training strategy in domain-adaptive pre-training are beneficial for the NER domain adaptation, and our proposed method can consistently outperform existing cross-domain NER baselines. Nevertheless, experiments also illustrate the challenge of this cross-domain NER task. We hope that our dataset and baselines will catalyze research in the NER domain adaptation area. The code and data are available at https://github.com/zliucr/CrossNER. Zihan Liu 0001, Yan Xu 0012, Tiezheng Yu, Wenliang Dai, Ziwei Ji 0001, Samuel Cahyawijaya, Andrea Madotto, Pascale Fung |
AAAI | 5 |