VLDB 2026 Research / reviewers in the wild / expert
Belinda Z. Li
dblp:263/9914
· DBLP profile ↗
12ranked-venue papers
8as first author
10since 2021 · last 2025
0000-0002-9961-180XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 8 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Eliciting Human Preferences with Language ModelsabstractLanguage models (LMs) can be directed to perform user- and context-dependent
tasks by using labeled examples or natural language prompts.
But selecting examples or writing prompts can be challenging---especially in tasks that require users to precisely articulate nebulous preferences or reason about complex edge cases. For such tasks, we introduce **Generative Active Task Elicitation (GATE)**, a method for using *LMs themselves* to guide the task specification process. GATE is a learning framework in which models elicit and infer human preferences through free-form, language-based interaction with users.
We identify prototypical challenges that users face when specifying preferences, and design three preference modeling tasks to study these challenges:
content recommendation, moral reasoning, and email validation.
In preregistered experiments, we show that LMs that learn to perform these tasks using GATE (by interactively querying users with open-ended questions) obtain preference specifications that are more informative than user-written prompts or examples. GATE matches existing task specification methods in the moral reasoning task, and significantly outperforms them in the content recommendation and email validation tasks. Users additionally report that interactive task elicitation requires less effort than prompting or example labeling and surfaces considerations that they did not anticipate on their own. Our findings suggest that LM-driven elicitation can be a powerful tool for aligning models to complex human preferences and values. Belinda Z. Li, Alex Tamkin, Noah D. Goodman, Jacob Andreas |
ICLR | 1 |
| 2025 | (How) Do Language Models Track State?abstractTransformer language models (LMs) exhibit behaviors—from storytelling to code generation—that seem to require tracking the unobserved state of an evolving world. How do they do this? We study state tracking in LMs trained or fine-tuned to compose permutations (i.e., to compute the order of a set of objects after a sequence of swaps). Despite the simple algebraic structure of this problem, many other tasks (e.g., simulation of finite automata and evaluation of boolean expressions) can be reduced to permutation composition, making it a natural model for state tracking in general. We show that LMs consistently learn one of two state tracking mechanisms for this task. The first closely resembles the “associative scan” construction used in recent theoretical work by Liu et al. (2023) and Merrill et al. (2024). The second uses an easy-to-compute feature (permutation parity) to partially prune the space of outputs, and then refines this with an associative scan. LMs that learn the former algorithm tend to generalize better and converge faster, and we show how to steer LMs toward one or the other with intermediate training tasks that encourage or suppress the heuristics. Our results demonstrate that transformer LMs, whether pre-trained or fine-tuned, can learn to implement efficient and interpretable state-tracking mechanisms, and the emergence of these mechanisms can be predicted and controlled. Code and data are available at https://github.com/belindal/state-tracking Belinda Z. Li, Zifan Carl Guo, Jacob Andreas |
ICML | 1 |
| 2025 | QuestBench: Can LLMs ask the right question to acquire information in reasoning tasks?abstractLarge language models (LLMs) have shown impressive performance on reasoning benchmarks like math and logic. While many works have largely assumed well-defined tasks, real-world queries are often underspecified and only solvable by acquiring missing information. We formalize this information-gathering problem as a constraint satisfaction problem (CSP) with missing variable assignments. Using a special case where only one necessary variable assignment is missing, we can evaluate an LLM's ability to identify the minimal necessary question to ask. We present QuestBench, a set of underspecified reasoning tasks solvable by asking at most one question, which includes: (1) Logic-Q: logical reasoning tasks with one missing proposition, (2) Planning-Q: PDDL planning problems with partially-observed initial states, (3) GSM-Q: human-annotated grade school math problems with one unknown variable, and (4) GSME-Q: equation-based version of GSM-Q. The LLM must select the correct clarification question from multiple options. While current models excel at GSM-Q and GSME-Q, they achieve only 40-50% accuracy on Logic-Q and Planning-Q. Analysis shows that the ability to solve well-specified reasoning problems is not sufficient for success on our benchmark: models struggle to identify the right question even when they can solve the fully specified version. This highlights the need for specifically optimizing models' information acquisition capabilities. Belinda Z. Li, Been Kim |
NeurIPS | 1 |
| 2024 | Preference-Conditioned Language-Guided AbstractionabstractLearning from demonstrations is a common way for users to teach robots, but it is prone to spurious feature correlations. Recent work constructs state abstractions, i.e. visual representations containing task-relevant features, from language as a way to perform more generalizable learning. However, these abstractions also depend on a user's preference for what matters in a task, which may be hard to describe or infeasible to exhaustively specify using language alone. How do we construct abstractions to capture these latent preferences? We observe that how humans behave reveals how they see the world. Our key insight is that changes in human behavior inform us that there are differences in preferences for how humans see the world, i.e. their state abstractions. In this work, we propose using language models (LMs) to query for those preferences directly given knowledge that a change in behavior has occurred. In our framework, we use the LM in two ways: first, given a text description of the task and knowledge of behavioral change between states, we query the LM for possible hidden preferences; second, given the most likely preference, we query the LM to construct the state abstraction. In this framework, the LM is also able to ask the human directly when uncertain about its own estimate. We demonstrate our framework's ability to construct effective preference-conditioned abstractions in simulated experiments, a user study, as well as on a real Spot robot performing mobile manipulation tasks. Andi Peng, Andreea Bobu, Belinda Z. Li, Theodore R. Sumers, Ilia Sucholutsky, Nishanth Kumar, Thomas L. Griffiths 0001, Julie A. Shah |
HRI | 3 |
| 2024 | Learning with Language-Guided State AbstractionsabstractWe describe a framework for using natural language to design state abstractions for imitation learning.
Generalizable policy learning in high-dimensional observation spaces is facilitated by well-designed state representations, which can surface important features of an environment and hide irrelevant ones.
These state representations are typically manually specified, or derived from other labor-intensive labeling procedures.
Our method, LGA (\textit{language-guided abstraction}), uses a combination of natural language supervision and background knowledge from language models (LMs) to automatically build state representations tailored to unseen tasks.
In LGA, a user first provides a (possibly incomplete) description of a target task in natural language; next, a pre-trained LM translates this task description into a state abstraction function that masks out irrelevant features; finally, an imitation policy is trained using a small number of demonstrations and LGA-generated abstract states.
Experiments on simulated robotic tasks show that LGA yields state abstractions similar to those designed by humans, but in a fraction of the time, and that these abstractions improve generalization and robustness in the presence of spurious correlations and ambiguous specifications.
We illustrate the utility of the learned abstractions on mobile manipulation tasks with a Spot robot. Andi Peng, Ilia Sucholutsky, Belinda Z. Li, Theodore R. Sumers, Thomas L. Griffiths 0001, Jacob Andreas, Julie A. Shah |
ICLR | 3 |
| 2023 | Toward Interactive DictationabstractVoice dictation is an increasingly important text input modality.Existing systems that allow both dictation and editing-by-voice restrict their command language to flat templates invoked by trigger words.In this work, we study the feasibility of allowing users to interrupt their dictation with spoken editing commands in open-ended natural language.We introduce a new task and dataset, TERTiUS, to experiment with such systems.To support this flexibility in real-time, a system must incrementally segment and classify spans of speech as either dictation or command, and interpret the spans that are commands.We experiment with using large pre-trained language models to predict the edited text, or alternatively, to predict a small text-editing program.Experiments show a natural trade-off between model accuracy and latency: a smaller model achieves 28% singlecommand interpretation accuracy with 1.3 seconds of latency, while a larger model achieves 55% with 7 seconds of latency. * Work performed during a research internship at Microsoft Semantic Machines.Just wanted to ask about the event on Friday the 23rd.Is the event still on?Just wanted to ask about the event on the 23rd, on Friday the 23rd.Is the event still on?Change"the event" to "it" in the last sentence.Just wanted to ask about the event on the 23rd.Just wanted to ask about the event on Friday the 23rd.Just wanted to check in about the event on Friday the 23rd.Is it still on? Belinda Z. Li, Jason Eisner, Adam Pauls, Sam Thomson |
ACL (1) | 1 |
| 2022 | Quantifying Adaptability in Pre-trained Language Models with 500 TasksabstractBelinda Li, Jane Yu, Madian Khabsa, Luke Zettlemoyer, Alon Halevy, Jacob Andreas. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Belinda Z. Li, Jane Dwivedi-Yu, Madian Khabsa, Luke Zettlemoyer, Alon Y. Halevy, Jacob Andreas |
NAACL-HLT | 1 |
| 2021 | Implicit Representations of Meaning in Neural Language ModelsabstractBelinda Z. Li, Maxwell Nye, Jacob Andreas. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Belinda Z. Li, Maxwell I. Nye, Jacob Andreas |
ACL/IJCNLP (1) | 1 |
| 2021 | On the Influence of Masking Policies in Intermediate Pre-trainingabstractCurrent NLP models are predominantly trained through a two-stage "pre-train then fine-tune" pipeline.Prior work has shown that inserting an intermediate pre-training stage, using heuristic masking policies for masked language modeling (MLM), can significantly improve final performance.However, it is still unclear (1) in what cases such intermediate pre-training is helpful, (2) whether hand-crafted heuristic objectives are optimal for a given task, and (3) whether a masking policy designed for one task is generalizable beyond that task.In this paper, we perform a large-scale empirical study to investigate the effect of various masking policies in intermediate pre-training with nine selected tasks across three categories.Crucially, we introduce methods to automate the discovery of optimal masking policies via direct supervision or meta-learning.We conclude that the success of intermediate pre-training is dependent on appropriate pre-train corpus, selection of output format (i.e., masked spans or full sentence), and clear understanding of the role that MLM plays for the downstream task.In addition, we find our learned masking policies outperform the heuristic of masking named entities on TriviaQA, and policies learned from one task can positively transfer to other tasks in certain cases, inviting future research in this direction. Qinyuan Ye, Belinda Z. Li, Sinong Wang, Benjamin Bolte, Hao Ma 0001, Scott Yih, Xiang Ren 0001, Madian Khabsa |
EMNLP (1) | 2 |
| 2021 | On Unifying Misinformation DetectionabstractNayeon Lee, Belinda Z. Li, Sinong Wang, Pascale Fung, Hao Ma, Wen-tau Yih, Madian Khabsa. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Nayeon Lee, Belinda Z. Li, Sinong Wang, Pascale Fung, Hao Ma 0001, Scott Yih, Madian Khabsa |
NAACL-HLT | 2 |
| 2020 | Active Learning for Coreference Resolution using Discrete AnnotationabstractWe improve upon pairwise annotation for active learning in coreference resolution, by asking annotators to identify mention antecedents if a presented mention pair is deemed not coreferent.This simple modification, when combined with a novel mention clustering algorithm for selecting which examples to label, is much more efficient in terms of the performance obtained per annotation budget.In experiments with existing benchmark coreference datasets, we show that the signal from this additional question leads to significant performance gains per human-annotation hour.Future work can use our annotation protocol to effectively develop coreference models for new domains.Our code is publicly available.1 Belinda Z. Li, Gabriel Stanovsky, Luke Zettlemoyer |
ACL | 1 |
| 2020 | Efficient One-Pass End-to-End Entity Linking for QuestionsabstractWe present ELQ, a fast end-to-end entity linking model for questions, which uses a biencoder to jointly perform mention detection and linking in one pass.Evaluated on WebQSP and GraphQuestions with extended annotations that cover multiple entities per question, ELQ outperforms the previous state of the art by a large margin of +12.7% and +19.6% F1, respectively.With a very fast inference time (1.57examples/s on a single CPU), ELQ can be useful for downstream question answering systems.In a proof-of-concept experiment, we demonstrate that using ELQ significantly improves the downstream QA performance of GraphRetriever (Min et al., 2019). 1 Belinda Z. Li, Sewon Min, Srinivasan Iyer 0001, Yashar Mehdad, Scott Yih |
EMNLP (1) | 1 |