Li Zhang 0039

dblp:89/5992-39 · also Li "Harry" Zhang · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
12since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 4 first-author · 12 since 2021Security and privacy · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Calibrating Large Language Models with Sample Consistency
abstract
Accurately gauging the confidence level of Large Language Models' (LLMs) predictions is pivotal for their reliable application. However, LLMs are often uncalibrated inherently and elude conventional calibration techniques due to their proprietary nature and massive scale. In this work, we derive model confidence from the distribution of multiple randomly sampled generations, using three measures of consistency. We extensively evaluate eleven open and closed-source models on nine reasoning datasets. Results show that consistency-based calibration methods outperform existing post-hoc approaches in terms of calibration error. Meanwhile, we find that factors such as intermediate explanations, model scaling, and larger sample sizes enhance calibration, while instruction-tuning makes calibration more difficult. Moreover, confidence scores obtained from consistency can potentially enhance model performance. Finally, we offer guidance on choosing suitable consistency metrics for calibration, tailored to model characteristics such as the exposure to instruction-tuning and RLHF.
Qing Lyu 0001, Kumar Shridhar, Chaitanya Malaviya, Li Zhang 0039, Yanai Elazar, Niket Tandon, Marianna Apidianaki, Mrinmaya Sachan, Chris Callison-Burch
AAAI4
2025 On the Limit of Language Models as Planning Formalizers
abstract
Large Language Models have been found to create plans that are neither executable nor verifiable in grounded environments.An emerging line of work demonstrates success in using the LLM as a formalizer to generate a formal representation of the planning domain in some language, such as Planning Domain Definition Language (PDDL).This formal representation can be deterministically solved to find a plan.We systematically evaluate this methodology while bridging some major gaps.While previous work only generates a partial PDDL representation, given templated, and therefore unrealistic environment descriptions, we generate the complete representation given descriptions of various naturalness levels.Among an array of observations critical to improve LLMs' formal planning abilities, we note that most large enough models can effectively formalize descriptions as PDDL, outperforming those directly generating plans, while being robust to lexical perturbation.As the descriptions become more natural-sounding, we observe a decrease in performance and provide detailed error analysis.1
Cassie Huang, Li Zhang 0039
ACL (1)2
2025 TurnaboutLLM: A Deductive Reasoning Benchmark from Detective Games
abstract
This paper introduces TURNABOUTLLM , a novel framework and dataset for evaluating the deductive reasoning abilities of Large Language Models (LLMs) by leveraging the interactive gameplay of detective games Ace Attorney and Danganronpa.The framework tasks LLMs with identifying contradictions between testimonies and evidences within long narrative contexts, a challenging task due to the large answer space and diverse reasoning types presented by its questions.We evaluate twelve state-of-the-art LLMs on the dataset, hinting at limitations of popular strategies for enhancing deductive reasoning such as extensive thinking and Chain-of-Thought prompting.The results also suggest varying effects of context size, the number of reasoning step and answer space size on model performance.Overall, TURN-ABOUTLLM presents a substantial challenge for LLMs' deductive reasoning abilities in complex, narrative-rich environments.1 * Equal contribution. 1 Our resources can be found at https://github.com/zharry29/ turnabout_llm.
Muyu He, Muhammad Adil Shahid, Li Zhang 0039
EMNLP6
2024 Choice-75: A Dataset on Decision Branching in Script Learning
abstract
Script learning studies how daily events unfold. It enables machines to reason about narratives with implicit information. Previous works mainly consider a script as a linear sequence of events while ignoring the potential branches that arise due to people’s circumstantial choices. We hence propose Choice-75, the first benchmark that challenges intelligent systems to make decisions given descriptive scenarios, containing 75 scripts and more than 600 scenarios. We also present preliminary results with current large language models (LLM). Although they demonstrate overall decent performances, there is still notable headroom in hard scenarios.
Zhaoyi Hou, Li Zhang 0039, Chris Callison-Burch
LREC/COLING2
2024 OpenPI2.0: An Improved Dataset for Entity Tracking in Texts
abstract
Li Zhang, Hainiu Xu, Abhinav Kommula, Chris Callison-Burch, Niket Tandon. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Li Zhang 0039, Hainiu Xu, Abhinav Kommula, Chris Callison-Burch, Niket Tandon
EACL (1)1
2023 Faithful Chain-of-Thought Reasoning
abstract
Qing Lyu, Shreya Havaldar, Adam Stein, Li Zhang, Delip Rao, Eric Wong, Marianna Apidianaki, Chris Callison-Burch. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Qing Lyu 0001, Shreya Havaldar, Adam Stein, Li Zhang 0039, Delip Rao, Eric Wong 0001, Marianna Apidianaki, Chris Callison-Burch
IJCNLP (1)4
2022 Show Me More Details: Discovering Hierarchies of Procedures from Semi-structured Web Data
abstract
Shuyan Zhou, Li Zhang, Yue Yang, Qing Lyu, Pengcheng Yin, Chris Callison-Burch, Graham Neubig. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Shuyan Zhou, Li Zhang 0039, Yue Yang 0006, Qing Lyu 0001, Chris Callison-Burch, Graham Neubig
ACL (1)2
2022 Unsupervised Entity Linking with Guided Summarization and Multiple-Choice Selection
abstract
Entity linking, the task of linking potentially ambiguous mentions in texts to corresponding knowledge-base entities, is an important component for language understanding.We address two challenge in entity linking: how to leverage wider contexts surrounding a mention, and how to deal with limited training data.We propose a fully unsupervised model called SumMC that first generates a guided summary of the contexts conditioning on the mention, and then casts the task to a multiple-choice problem where the model chooses an entity from a list of candidates.In addition to evaluating our model on existing datasets that focus on named entities, we create a new dataset that links noun phrases from WikiHow to Wikidata.We show that our SumMC model achieves stateof-the-art unsupervised performance on our new dataset and on existing datasets.
Li Zhang 0039, Chris Callison-Burch
EMNLP2
2022 Is "My Favorite New Movie" My Favorite Movie? Probing the Understanding of Recursive Noun Phrases
abstract
Qing Lyu, Zheng Hua, Daoxin Li, Li Zhang, Marianna Apidianaki, Chris Callison-Burch. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Qing Lyu 0001, Daoxin Li, Li Zhang 0039, Marianna Apidianaki, Chris Callison-Burch
NAACL-HLT4
2022 Label Definitions Improve Semantic Role Labeling
abstract
Argument classification is at the core of Semantic Role Labeling.Given a sentence and the predicate, a semantic role label is assigned to each argument of the predicate.While semantic roles come with meaningful definitions, existing work has treated them as symbolic.Learning symbolic labels usually requires ample training data, which is frequently unavailable due to the cost of annotation.We instead propose to retrieve and leverage the definitions of these labels from the annotation guidelines.For example, the verb predicate "work" has arguments defined as "worker", "job", "employer", etc.Our model achieves state-of-theart performance on the CoNLL09 English SRL dataset injected with label definitions given the predicate senses.The performance improvement is even more pronounced in low-resource settings when training data is scarce. 1
Li Zhang 0039, Ishan Jindal, Yunyao Li 0001
NAACL-HLT1
2021 Visual Goal-Step Inference using wikiHow
abstract
Understanding what sequence of steps are needed to complete a goal can help artificial intelligence systems reason about human activities.Past work in NLP has examined the task of goal-step inference for text.We introduce the visual analogue.We propose the Visual Goal-Step Inference (VGSI) task, where a model is given a textual goal and must choose which of four images represents a plausible step towards that goal.With a new dataset harvested from wikiHow consisting of 772,277 images representing human actions, we show that our task is challenging for state-of-theart multimodal models.Moreover, the multimodal representation learned from our data can be effectively transferred to other datasets like HowTo100m, increasing the VGSI accuracy by 15 -20%.Our task will facilitate multimodal reasoning about procedural events.
Yue Yang 0006, Artemis Panagopoulou, Qing Lyu 0001, Li Zhang 0039, Mark Yatskar, Chris Callison-Burch
EMNLP (1)4
2021 Goal-Oriented Script Construction
abstract
The knowledge of scripts, common chains of events in stereotypical scenarios, is a valuable asset for task-oriented natural language understanding systems.We propose the Goal-Oriented Script Construction task, where a model produces a sequence of steps to accomplish a given goal.We pilot our task on the first multilingual script learning dataset supporting 18 languages collected from wikiHow, a website containing half a million how-to articles.For baselines, we consider both a generationbased approach using a language model and a retrieval-based approach by first retrieving the relevant steps from a large candidate pool and then ordering them.We show that our task is practical, feasible but challenging for state-of-the-art Transformer models, and that our methods can be readily deployed for various other datasets and domains with decent zero-shot performance 1 . * Equal contribution.
Qing Lyu 0001, Li Zhang 0039, Chris Callison-Burch
INLG2
2020 Reasoning about Goals, Steps, and Temporal Ordering with WikiHow
abstract
We propose a suite of reasoning tasks on two types of relations between procedural events: goal-step relations ("learn poses" is a step in the larger goal of "doing yoga") and step-step temporal relations ("buy a yoga mat" typically precedes "learn poses"). We introduce a dataset targeting these two relations based on wikiHow, a website of instructional how-to articles. Our human-validated test set serves as a reliable benchmark for commonsense inference, with a gap of about 10% to 20% between the performance of state-of-the-art transformer models and human performance. Our automatically-generated training set allows models to effectively transfer to out-of-domain tasks requiring knowledge of procedural events, with greatly improved performances on SWAG, Snips, and the Story Cloze Test in zero- and few-shot settings.
Li Zhang 0039, Qing Lyu 0001, Chris Callison-Burch
EMNLP (1)1
2020 Small but Mighty: New Benchmarks for Split and Rephrase
abstract
Split and Rephrase is a text simplification task of rewriting a complex sentence into simpler ones.As a relatively new task, it is paramount to ensure the soundness of its evaluation benchmark and metric.We find that the widely used benchmark dataset universally contains easily exploitable syntactic cues caused by its automatic generation process.Taking advantage of such cues, we show that even a simple rule-based model can perform on par with the state-of-the-art model.To remedy such limitations, we collect and release two crowdsourced benchmark datasets.We not only make sure that they contain significantly more diverse syntax, but also carefully control for their quality according to a welldefined set of criteria.While no satisfactory automatic metric exists, we apply fine-grained manual evaluation based on these criteria using crowdsourcing, showing that our datasets better represent the task and are significantly more challenging for the models. 1
Li Zhang 0039, Huaiyu Zhu 0001, Siddhartha Brahma, Yunyao Li 0001
EMNLP (1)1
2019 Kindness is a Risky Business: On the Usage of the Accessibility APIs in Android
Wenrui Diao, Yue Zhang 0025, Li Zhang 0039, Zhou Li 0001, Fenghao Xu, Xiaorui Pan, Jian Weng 0001, Kehuan Zhang, XiaoFeng Wang 0001
RAID3
2019 CryptoREX: Large-scale Analysis of Cryptographic Misuse in IoT Devices
Li Zhang 0039, Jiongyi Chen, Wenrui Diao, Shanqing Guo, Jian Weng 0001, Kehuan Zhang
RAID1
2018 Improving Text-to-SQL Evaluation Methodology
abstract
Catherine Finegan-Dollak, Jonathan K. Kummerfeld, Li Zhang, Karthik Ramanathan, Sesh Sadasivam, Rui Zhang, Dragomir Radev. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Catherine Finegan-Dollak, Jonathan K. Kummerfeld, Li Zhang 0039, Karthik Ramanathan, Sesh Sadasivam, Rui Zhang 0037, Dragomir R. Radev
ACL (1)3