Caiqi Zhang

dblp:301/5843 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
13since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 first-author · 11 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Value of Information: A Framework for Human-Agent Communication
abstract
Yijiang River Dong, Tiancheng Hu, Zheng Hui, Caiqi Zhang, Ivan Vulić, Andreea Bobu, Nigel Collier. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yijiang River Dong, Tiancheng Hu, Zheng Hui, Caiqi Zhang, Ivan Vulic, Andreea Bobu, Nigel Collier
ACL (1)4
2026 LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form Generation
abstract
Hallucination remains a major challenge for the safe and trustworthy deployment of large language models (LLMs) in factual content generation.Prior work has explored confidence estimation as an effective approach to hallucination detection, but often relies on post-hoc self-consistency methods that require computationally expensive sampling.Verbalized confidence offers a more efficient alternative, but existing approaches are largely limited to shortform question answering (QA) tasks and do not generalize well to open-ended generation.In this paper, we propose LOVEC (Long-form Verbalized Confidence), a novel reinforcement learning (RL)-based method that trains LLMs to append an on-the-fly numerical confidence score to each generated statement during longform generation.The confidence score serves as a direct and interpretable signal of the factuality of generation.We introduce two evaluation settings, free-form tagging and iterative tagging, to assess different verbalized confidence estimation methods.Experiments on three long-form QA datasets show that our RLtrained models achieve better calibration and generalize robustly across domains.Also, our method is highly efficient, being 20× faster than traditional self-consistency methods while achieving better calibration.
Caiqi Zhang, Chengzu Li, Nigel Collier, Andreas Vlachos 0001
ACL (1)1
2025 LoGU: Long-form Generation with Uncertainty Expressions
abstract
Ruihan Yang, Caiqi Zhang, Zhisong Zhang, Xinting Huang, Sen Yang, Nigel Collier, Dong Yu, Deqing Yang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Ruihan Yang, Caiqi Zhang, Zhisong Zhang, Xinting Huang, Nigel Collier, Dong Yu 0001, Deqing Yang
ACL (1)2
2025 Conformity in Large Language Models
abstract
The conformity effect describes the tendency of individuals to align their responses with the majority.Studying this bias in large language models (LLMs) is crucial, as LLMs are increasingly used in various information-seeking and decision-making tasks as conversation partners to improve productivity.Thus, conformity to incorrect responses can compromise their effectiveness.In this paper, we adapt psychological experiments to examine the extent of conformity in popular LLMs.Our findings reveal that all tested models exhibit varying levels of conformity toward the majority, regardless of their initial choice or correctness, across different knowledge domains.Notably, we are the first to show that LLMs are more likely to conform when they are more uncertain in their own prediction.We further explore factors that influence conformity, such as training paradigms and input characteristics, finding that instruction-tuned models are less susceptible to conformity, while increasing the naturalness of majority tones amplifies conformity.Finally, we propose two interventions, Devil's Advocate and Question Distillation, to mitigate conformity, providing insights into building more robust language models. What is the oldest college in Cambridge?It is Peterhouse College. What is the oldest college in Cambridge?King's.
Caiqi Zhang, Tom Stafford 0002, Nigel Collier, Andreas Vlachos 0001
ACL (1)2
2025 A Head to Predict and a Head to Question: Pre-trained Uncertainty Quantification Heads for Hallucination Detection in LLM Outputs
abstract
Artem Shelmanov, Ekaterina Fadeeva, Akim Tsvigun, Ivan Tsvigun, Zhuohan Xie, Igor Kiselev, Nico Daheim, Caiqi Zhang, Artem Vazhentsev, Mrinmaya Sachan, Preslav Nakov, Timothy Baldwin. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Artem Shelmanov, Ekaterina Fadeeva, Akim Tsvigun, Ivan Tsvigun, Zhuohan Xie, Igor Kiselev, Nico Daheim, Caiqi Zhang, Artem Vazhentsev, Mrinmaya Sachan, Preslav Nakov, Timothy Baldwin
EMNLP8
2025 UNCLE: Benchmarking Uncertainty Expressions in Long-Form Generation
abstract
Large Language Models (LLMs) are prone to hallucination, particularly in long-form generations.A promising direction to mitigate hallucination is to teach LLMs to express uncertainty explicitly when they lack sufficient knowledge.However, existing work lacks direct and fair evaluation of LLMs' ability to express uncertainty effectively in long-form generation.To address this gap, we first introduce UNCLE, a benchmark designed to evaluate uncertainty expression in both long-and short-form question answering (QA).UNCLE covers five domains and includes more than 1,000 entities, each with paired short-and long-form QA items.Our dataset is the first to directly link short-and long-form QA through aligned questions and gold-standard answers.Along with UNCLE, we propose a suite of new metrics to assess the models' capabilities to selectively express uncertainty.We then demonstrate that current models fail to convey uncertainty appropriately in long-form generation.We further explore both prompt-based and training-based methods to improve models' performance, with the training-based methods yielding greater gains.Further analysis of alignment gaps between short-and long-form uncertainty expression highlights promising directions for future research using UNCLE.
Ruihan Yang, Caiqi Zhang, Zhisong Zhang, Xinting Huang, Dong Yu 0001, Nigel Collier, Deqing Yang
EMNLP2
2025 All Roads Lead to Rome: Graph-Based Confidence Estimation for Large Language Model Reasoning
abstract
Confidence estimation is essential for the reliable deployment of large language models (LLMs).Existing methods are primarily designed for factual QA tasks and often fail to generalize to reasoning tasks.To address this gap, we propose a set of training-free, graph-based confidence estimation methods tailored to reasoning tasks.Our approach models reasoning paths as directed graphs and estimates confidence by exploiting graph properties such as centrality, path convergence, and path weighting.Experiments with two LLMs on three reasoning datasets demonstrate improved confidence estimation and enhanced performance on two downstream tasks.
Caiqi Zhang, Ehsan Shareghi, Nigel Collier
EMNLP1
2024 TopViewRS: Vision-Language Models as Top-View Spatial Reasoners
abstract
Top-view perspective denotes a typical way in which humans read and reason over different types of maps, and it is vital for localization and navigation of humans as well as of 'non-human' agents, such as the ones backed by large Vision-Language Models (VLMs).Nonetheless, spatial reasoning capabilities of modern VLMs in this setup remain unattested and underexplored.In this work, we study their capability to understand and reason over spatial relations from the top view.The focus on top view also enables controlled evaluations at different granularity of spatial reasoning; we clearly disentangle different abilities (e.g., recognizing particular objects versus understanding their relative positions).We introduce the TOPVIEWRS (Top-View Reasoning in Space) dataset, consisting of 11,384 multiple-choice questions with either realistic or semantic top-view map as visual input.We then use it to study and evaluate VLMs across 4 perception and reasoning tasks with different levels of complexity.Evaluation of 10 representative open-and closedsource VLMs reveals the gap of more than 50% compared to average human performance, and it is even lower than the random baseline in some cases.Although additional experiments show that Chain-of-Thought reasoning can boost model capabilities by 5.82% on average, the overall performance of VLMs remains limited.Our findings underscore the critical need for enhanced model capability in top-view spatial reasoning and set a foundation for further research towards human-level proficiency of VLMs in real-world multimodal tasks.
Chengzu Li, Caiqi Zhang, Han Zhou 0010, Nigel Collier, Anna Korhonen, Ivan Vulic
EMNLP2
2024 LUQ: Long-text Uncertainty Quantification for LLMs
abstract
Large Language Models (LLMs) have demonstrated remarkable capability in a variety of NLP tasks.However, LLMs are also prone to generate nonfactual content.Uncertainty Quantification (UQ) is pivotal in enhancing our understanding of a model's confidence on its generation, thereby aiding in the mitigation of nonfactual outputs.Existing research on UQ predominantly targets short text generation, typically yielding brief, word-limited responses.However, real-world applications frequently necessitate much longer responses.Our study first highlights the limitations of current UQ methods in handling long text generation.We then introduce LUQ with its two variations: LUQ-ATOMIC and LUQ-PAIR, a series of novel sampling-based UQ approaches specifically designed for long text.Our findings reveal that LUQ outperforms existing baseline methods in correlating with the model's factuality scores (negative coefficient of -0.85 observed for Gemini Pro).To further improve the factuality of LLM responses, we propose LUQ-ENSEMBLE, a method that ensembles responses from multiple models and selects the response with the lowest uncertainty.The ensembling method greatly improves the response factuality upon the best standalone LLM. 1 * Now at Google DeepMind.
Caiqi Zhang, Fangyu Liu 0001, Marco Basaldella, Nigel Collier
EMNLP1
2024 Do We Need Language-Specific Fact-Checking Models? The Case of Chinese
abstract
This paper investigates the potential benefits of language-specific fact-checking models, focusing on the case of Chinese using CHEF dataset.To better reflect real-world fact-checking, we first develop a novel Chinese document-level evidence retriever, achieving state-of-the-art performance.We then demonstrate the limitations of translation-based methods and multilingual language models, highlighting the need for language-specific systems.To better analyze token-level biases in different systems, we construct an adversarial dataset based on the CHEF dataset, where each instance has a large word overlap with the original one but holds the opposite veracity label.Experimental results on the CHEF dataset and our adversarial dataset show that our proposed method outperforms translation-based methods and multilingual language models and is more robust toward biases, emphasizing the importance of language-specific fact-checking systems. 1 Verifiers Retrievers
Caiqi Zhang, Zhijiang Guo, Andreas Vlachos 0001
EMNLP1
2023 Learning Action Conditions from Instructional Manuals for Instruction Understanding
abstract
The ability to infer pre-and postconditions of an action is vital for comprehending complex instructions, and is essential for applications such as autonomous instruction-guided agents and assistive AI that supports humans to perform physical tasks.In this work, we propose a task dubbed action condition inference, which extracts mentions of preconditions and postconditions of actions in instructional manuals.We propose a weakly supervised approach utilizing automatically constructed large-scale training instances from online instructions, and curate a densely human-annotated and validated dataset to study how well the current NLP models do on the proposed task.We design two types of models differ by whether contextualized and global information is leveraged, as well as various combinations of heuristics to construct the weak supervisions.Our experiments show a >20% F1-score improvement with considering the entire instruction contexts and a > 6% F1-score benefit with the proposed heuristics.However, the best performing model is still well-behind human performance.1 standalone Heuristics Examples Descriptions Entity-Tracing & Coref.… Slice 500 grams of onions.… … Heat the pan with olive oil.… … Place them in the frying pan.… Precondition 1 Precondition 2The shared entities are pan and onions (linked via co-references to them).
Te-Lin Wu, Caiqi Zhang, Alexander Spangher, Nanyun Peng 0001
ACL (1)2
2023 Hybrid Learning for Mobile Ad-Hoc Distancing/Positioning Using Bluetooth Low Energy
abstract
With the advent of Bluetooth low-energy (BLE)-enabled smartphones, there has been considerable interest in investigating BLE-based distancing/positioning methods (e.g., for social distancing applications). In this article, we present a novel hybrid learning method to support mobile ad-hoc distancing (MAD)/positioning (MAP) using BLE-enabled smartphones. Compared to traditional BLE-based distancing/positioning methods, the hybrid learning method provides the following unique features and contributions. First, it combines unsupervised learning, supervised learning, and genetic algorithms (GAs) for enhancing distance estimation accuracy. Second, unsupervised learning is employed to identify three pseudo channels/clusters for enhanced RSSI data processing. Third, its underlying mechanism is based on a new pattern-inspired approach to enhance the machine learning process. Fourth, it provides a flagging mechanism to alert users if a predicted distance is accurate or not. Fifth, it provides a model aggregation scheme with an innovative 2-D GA to aggregate the distance estimation results of different machine learning models. As an application of hybrid learning for distance estimation, we also present a new MAP scenario with an iterative algorithm to estimate mobile positions in an ad-hoc environment. Experimental results show the effectiveness of the hybrid learning method. In particular, hybrid learning without flagging and with flagging outperforms the baseline by 57% and 65%, respectively, in terms of mean absolute error. By means of model aggregation, a further 4% improvement can be realized. The hybrid learning approach can also be applied to previous work to enhance distance estimation accuracy and provide valuable insights for further research.
Yik Him Ho, Caiqi Zhang, Yerkezhan Sartayeva, Henry C. B. Chan
IEEE Internet Things J.3
2021 PRUID: Practical User Interface Distribution for Multi-surface Computing
abstract
It becomes more and more common for people to have multiple mobile devices. This opens the opportunity of multi-surface computing in which users interact with an app using multiple devices simultaneously. Recently, a system called FLUID was developed, which can distribute User Interface (UI) elements of an app to multiple devices to support multi-surface computing. FLUID enables general, flexible and transparent multi-device interaction, which cannot be achieved by previous approaches such as screen mirroring, app migration, and customized app development on multiple devices. However, the practicality of FLUID is still severely limited because it requires that (1) the app source codes must be available and (2) the same app is pre-installed on all devices. This paper presents PRUID, a UI distribution system that is free from the above-mentioned limitations of FLUID. PRUID captures and extracts relevant information about UI elements to be distributed completely at run time, without requiring the app source code. An app-independent UI agent is designed to dock and render the UI components distributed to the guest device, so pre-installation of the app on guest devices is not required. We developed representative use cases to demonstrate the usage and evaluate the performance of PRUID. The evaluation results show that the extra overhead incurred due to the UI information extraction at run time is marginal and PRUID provides a smooth user experience.
Menglong Cui, Mingsong Lv, Qingqiang He, Caiqi Zhang, Chuancai Gu, Tao Yang 0024, Nan Guan
DAC4