VLDB 2026 Research / reviewers in the wild / expert
Li Zhang 0039
dblp:89/5992-39 · also Li "Harry" Zhang
· DBLP profile ↗
17ranked-venue papers
5as first author
12since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 4 first-author · 12 since 2021Security and privacy · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Calibrating Large Language Models with Sample ConsistencyabstractAccurately gauging the confidence level of Large Language Models' (LLMs) predictions is pivotal for their reliable application. However, LLMs are often uncalibrated inherently and elude conventional calibration techniques due to their proprietary nature and massive scale. In this work, we derive model confidence from the distribution of multiple randomly sampled generations, using three measures of consistency. We extensively evaluate eleven open and closed-source models on nine reasoning datasets. Results show that consistency-based calibration methods outperform existing post-hoc approaches in terms of calibration error. Meanwhile, we find that factors such as intermediate explanations, model scaling, and larger sample sizes enhance calibration, while instruction-tuning makes calibration more difficult. Moreover, confidence scores obtained from consistency can potentially enhance model performance. Finally, we offer guidance on choosing suitable consistency metrics for calibration, tailored to model characteristics such as the exposure to instruction-tuning and RLHF. Qing Lyu 0001, Kumar Shridhar, Chaitanya Malaviya, Li Zhang 0039, Yanai Elazar, Niket Tandon, Marianna Apidianaki, Mrinmaya Sachan, Chris Callison-Burch |
AAAI | 4 |
| 2025 | On the Limit of Language Models as Planning FormalizersabstractLarge Language Models have been found to create plans that are neither executable nor verifiable in grounded environments.An emerging line of work demonstrates success in using the LLM as a formalizer to generate a formal representation of the planning domain in some language, such as Planning Domain Definition Language (PDDL).This formal representation can be deterministically solved to find a plan.We systematically evaluate this methodology while bridging some major gaps.While previous work only generates a partial PDDL representation, given templated, and therefore unrealistic environment descriptions, we generate the complete representation given descriptions of various naturalness levels.Among an array of observations critical to improve LLMs' formal planning abilities, we note that most large enough models can effectively formalize descriptions as PDDL, outperforming those directly generating plans, while being robust to lexical perturbation.As the descriptions become more natural-sounding, we observe a decrease in performance and provide detailed error analysis.1 Cassie Huang, Li Zhang 0039 |
ACL (1) | 2 |
| 2025 | TurnaboutLLM: A Deductive Reasoning Benchmark from Detective GamesabstractThis paper introduces TURNABOUTLLM , a novel framework and dataset for evaluating the deductive reasoning abilities of Large Language Models (LLMs) by leveraging the interactive gameplay of detective games Ace Attorney and Danganronpa.The framework tasks LLMs with identifying contradictions between testimonies and evidences within long narrative contexts, a challenging task due to the large answer space and diverse reasoning types presented by its questions.We evaluate twelve state-of-the-art LLMs on the dataset, hinting at limitations of popular strategies for enhancing deductive reasoning such as extensive thinking and Chain-of-Thought prompting.The results also suggest varying effects of context size, the number of reasoning step and answer space size on model performance.Overall, TURN-ABOUTLLM presents a substantial challenge for LLMs' deductive reasoning abilities in complex, narrative-rich environments.1 * Equal contribution. 1 Our resources can be found at https://github.com/zharry29/ turnabout_llm. Muyu He, Muhammad Adil Shahid, Li Zhang 0039 |
EMNLP | 6 |
| 2024 | Choice-75: A Dataset on Decision Branching in Script LearningabstractScript learning studies how daily events unfold. It enables machines to reason about narratives with implicit information. Previous works mainly consider a script as a linear sequence of events while ignoring the potential branches that arise due to people’s circumstantial choices. We hence propose Choice-75, the first benchmark that challenges intelligent systems to make decisions given descriptive scenarios, containing 75 scripts and more than 600 scenarios. We also present preliminary results with current large language models (LLM). Although they demonstrate overall decent performances, there is still notable headroom in hard scenarios. Zhaoyi Hou, Li Zhang 0039, Chris Callison-Burch |
LREC/COLING | 2 |
| 2024 | OpenPI2.0: An Improved Dataset for Entity Tracking in TextsabstractLi Zhang, Hainiu Xu, Abhinav Kommula, Chris Callison-Burch, Niket Tandon. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Li Zhang 0039, Hainiu Xu, Abhinav Kommula, Chris Callison-Burch, Niket Tandon |
EACL (1) | 1 |
| 2023 | Faithful Chain-of-Thought ReasoningabstractQing Lyu, Shreya Havaldar, Adam Stein, Li Zhang, Delip Rao, Eric Wong, Marianna Apidianaki, Chris Callison-Burch. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Qing Lyu 0001, Shreya Havaldar, Adam Stein, Li Zhang 0039, Delip Rao, Eric Wong 0001, Marianna Apidianaki, Chris Callison-Burch |
IJCNLP (1) | 4 |
| 2022 | Show Me More Details: Discovering Hierarchies of Procedures from Semi-structured Web DataabstractShuyan Zhou, Li Zhang, Yue Yang, Qing Lyu, Pengcheng Yin, Chris Callison-Burch, Graham Neubig. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Shuyan Zhou, Li Zhang 0039, Yue Yang 0006, Qing Lyu 0001, Chris Callison-Burch, Graham Neubig |
ACL (1) | 2 |
| 2022 | Unsupervised Entity Linking with Guided Summarization and Multiple-Choice SelectionabstractEntity linking, the task of linking potentially ambiguous mentions in texts to corresponding knowledge-base entities, is an important component for language understanding.We address two challenge in entity linking: how to leverage wider contexts surrounding a mention, and how to deal with limited training data.We propose a fully unsupervised model called SumMC that first generates a guided summary of the contexts conditioning on the mention, and then casts the task to a multiple-choice problem where the model chooses an entity from a list of candidates.In addition to evaluating our model on existing datasets that focus on named entities, we create a new dataset that links noun phrases from WikiHow to Wikidata.We show that our SumMC model achieves stateof-the-art unsupervised performance on our new dataset and on existing datasets. Li Zhang 0039, Chris Callison-Burch |
EMNLP | 2 |
| 2022 | Is "My Favorite New Movie" My Favorite Movie? Probing the Understanding of Recursive Noun PhrasesabstractQing Lyu, Zheng Hua, Daoxin Li, Li Zhang, Marianna Apidianaki, Chris Callison-Burch. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Qing Lyu 0001, Daoxin Li, Li Zhang 0039, Marianna Apidianaki, Chris Callison-Burch |
NAACL-HLT | 4 |
| 2022 | Label Definitions Improve Semantic Role LabelingabstractArgument classification is at the core of Semantic Role Labeling.Given a sentence and the predicate, a semantic role label is assigned to each argument of the predicate.While semantic roles come with meaningful definitions, existing work has treated them as symbolic.Learning symbolic labels usually requires ample training data, which is frequently unavailable due to the cost of annotation.We instead propose to retrieve and leverage the definitions of these labels from the annotation guidelines.For example, the verb predicate "work" has arguments defined as "worker", "job", "employer", etc.Our model achieves state-of-theart performance on the CoNLL09 English SRL dataset injected with label definitions given the predicate senses.The performance improvement is even more pronounced in low-resource settings when training data is scarce. 1 Li Zhang 0039, Ishan Jindal, Yunyao Li 0001 |
NAACL-HLT | 1 |
| 2021 | Visual Goal-Step Inference using wikiHowabstractUnderstanding what sequence of steps are needed to complete a goal can help artificial intelligence systems reason about human activities.Past work in NLP has examined the task of goal-step inference for text.We introduce the visual analogue.We propose the Visual Goal-Step Inference (VGSI) task, where a model is given a textual goal and must choose which of four images represents a plausible step towards that goal.With a new dataset harvested from wikiHow consisting of 772,277 images representing human actions, we show that our task is challenging for state-of-theart multimodal models.Moreover, the multimodal representation learned from our data can be effectively transferred to other datasets like HowTo100m, increasing the VGSI accuracy by 15 -20%.Our task will facilitate multimodal reasoning about procedural events. Yue Yang 0006, Artemis Panagopoulou, Qing Lyu 0001, Li Zhang 0039, Mark Yatskar, Chris Callison-Burch |
EMNLP (1) | 4 |
| 2021 | Goal-Oriented Script ConstructionabstractThe knowledge of scripts, common chains of events in stereotypical scenarios, is a valuable asset for task-oriented natural language understanding systems.We propose the Goal-Oriented Script Construction task, where a model produces a sequence of steps to accomplish a given goal.We pilot our task on the first multilingual script learning dataset supporting 18 languages collected from wikiHow, a website containing half a million how-to articles.For baselines, we consider both a generationbased approach using a language model and a retrieval-based approach by first retrieving the relevant steps from a large candidate pool and then ordering them.We show that our task is practical, feasible but challenging for state-of-the-art Transformer models, and that our methods can be readily deployed for various other datasets and domains with decent zero-shot performance 1 . * Equal contribution. Qing Lyu 0001, Li Zhang 0039, Chris Callison-Burch |
INLG | 2 |
| 2020 | Reasoning about Goals, Steps, and Temporal Ordering with WikiHowabstractWe propose a suite of reasoning tasks on two types of relations between procedural events: goal-step relations ("learn poses" is a step in the larger goal of "doing yoga") and step-step temporal relations ("buy a yoga mat" typically precedes "learn poses"). We introduce a dataset targeting these two relations based on wikiHow, a website of instructional how-to articles. Our human-validated test set serves as a reliable benchmark for commonsense inference, with a gap of about 10% to 20% between the performance of state-of-the-art transformer models and human performance. Our automatically-generated training set allows models to effectively transfer to out-of-domain tasks requiring knowledge of procedural events, with greatly improved performances on SWAG, Snips, and the Story Cloze Test in zero- and few-shot settings. Li Zhang 0039, Qing Lyu 0001, Chris Callison-Burch |
EMNLP (1) | 1 |
| 2020 | Small but Mighty: New Benchmarks for Split and RephraseabstractSplit and Rephrase is a text simplification task of rewriting a complex sentence into simpler ones.As a relatively new task, it is paramount to ensure the soundness of its evaluation benchmark and metric.We find that the widely used benchmark dataset universally contains easily exploitable syntactic cues caused by its automatic generation process.Taking advantage of such cues, we show that even a simple rule-based model can perform on par with the state-of-the-art model.To remedy such limitations, we collect and release two crowdsourced benchmark datasets.We not only make sure that they contain significantly more diverse syntax, but also carefully control for their quality according to a welldefined set of criteria.While no satisfactory automatic metric exists, we apply fine-grained manual evaluation based on these criteria using crowdsourcing, showing that our datasets better represent the task and are significantly more challenging for the models. 1 Li Zhang 0039, Huaiyu Zhu 0001, Siddhartha Brahma, Yunyao Li 0001 |
EMNLP (1) | 1 |
| 2019 | Kindness is a Risky Business: On the Usage of the Accessibility APIs in Android
Wenrui Diao, Yue Zhang 0025, Li Zhang 0039, Zhou Li 0001, Fenghao Xu, Xiaorui Pan, Jian Weng 0001, Kehuan Zhang, XiaoFeng Wang 0001 |
RAID | 3 |
| 2019 | CryptoREX: Large-scale Analysis of Cryptographic Misuse in IoT Devices
Li Zhang 0039, Jiongyi Chen, Wenrui Diao, Shanqing Guo, Jian Weng 0001, Kehuan Zhang |
RAID | 1 |
| 2018 | Improving Text-to-SQL Evaluation MethodologyabstractCatherine Finegan-Dollak, Jonathan K. Kummerfeld, Li Zhang, Karthik Ramanathan, Sesh Sadasivam, Rui Zhang, Dragomir Radev. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Catherine Finegan-Dollak, Jonathan K. Kummerfeld, Li Zhang 0039, Karthik Ramanathan, Sesh Sadasivam, Rui Zhang 0037, Dragomir R. Radev |
ACL (1) | 3 |