Ben Zhou

dblp:219/5276 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 6 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Toward Controllable and Trustworthy LLM Reasoning: From Failure Mapping to Cognition-inspired Control and Real-world Impact
abstract
Large Language Models (LLMs) have advanced rapidly and raised the bar for what AI is expected to do. However, accompanied with such progress is a stronger consensus that these models consistently fail in out-of-distribution reasoning, especially on tasks that require abstraction, transfer, or long-horizon planning. While acceptable for most consumer use, these issues prevent AI from being safely deployed in high-stakes settings (e.g., healthcare), where stakeholders cannot trust AI models that exhibit uncontrollable and unpredictable failures. In this talk, I will discuss our work and insights on how to make LLM reasoning controllable and trustworthy, by 1) understanding the mechanisms of LLM reasoning and predicting when LLM will fail; 2) improving model reasoning and generalization based on such insights; and 3) moving towards trustworthy AI applications through such improvements, and identifying new problems to form a healthy positive-feedback loop.
Ben Zhou
AAAI1
2025 ThinkTuning: Instilling Cognitive Reflections without Distillation
abstract
Recent advances in test-time scaling have led to the emergence of thinking LLMs that exhibit self-reflective behaviors and multi-step reasoning.While RL drives this self-improvement paradigm, a recent study (Gandhi et al., 2025) shows that RL alone does not truly instill these new reasoning abilities -it merely draws out behaviors already present in the base models.This raises a question: How can we train models that don't exhibit such thinking behavior to develop it in the first place?To this end, we propose THINKTUNING, a GRPO-based interactive training approach where we augment the rollouts of a student model with the guidance from a teacher model.A simple idea from classroom practice inspires our method: a teacher poses a problem, lets the student try an answer, then gives corrective feedback-enough to point the mind in the right direction and then show the solution.Each piece of feedback reshapes the student's thoughts, leading them to arrive at the correct solution.Similarly, we find that this type of implicit supervision through feedback from a teacher model of the same size improves the reasoning capabilities of the student model.In particular, on average, our method shows a 3.85% improvement over zero-shot baselines across benchmarks, and on MATH-500, AIME and GPQA-Diamond it shows 2.08%, 2.23% and 3.99% improvements over the vanilla-GRPO baseline 1 .
Aswin RRV, Jacob Dineen, Divij Handa, Md Nayem Uddin, Mihir Parmar, Chitta Baral, Ben Zhou
EMNLP7
2025 BIRD: A Trustworthy Bayesian Inference Framework for Large Language Models
abstract
Predictive models often need to work with incomplete information in real-world tasks. Consequently, they must provide reliable probability or confidence estimation, especially in large-scale decision-making and planning tasks. Current large language models (LLMs) are insufficient for accurate estimations, but they can generate relevant factors that may affect the probabilities, produce coarse-grained probabilities when the information is more complete, and help determine which factors are relevant to specific downstream contexts. In this paper, we make use of these capabilities of LLMs to provide a significantly more accurate probabilistic estimation. We propose BIRD, a novel probabilistic inference framework that aligns a Bayesian network with LLM abductions and then estimates more accurate probabilities in a deduction step. We show BIRD provides reliable probability estimations that are 30% better than those provided directly by LLM baselines. These estimates further contribute to better and more trustworthy decision making.
Yu Feng 0013, Ben Zhou, Dan Roth 0001
ICLR2
2025 ToW: Thoughts of Words Improve Reasoning in Large Language Models
abstract
Zhikun Xu, Ming Shen, Jacob Dineen, Zhaonan Li, Xiao Ye, Shijie Lu, Aswin Rrv, Chitta Baral, Ben Zhou. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Zhikun Xu, Ming Shen 0006, Jacob Dineen, Shijie Lu, Aswin RRV, Chitta Baral, Ben Zhou
NAACL (Long Papers)9
2024 Sub-Sentence Encoder: Contrastive Learning of Propositional Semantic Representations
abstract
Sihao Chen, Hongming Zhang, Tong Chen, Ben Zhou, Wenhao Yu, Dian Yu, Baolin Peng, Hongwei Wang, Dan Roth, Dong Yu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Hongming Zhang 0009, Ben Zhou, Wenhao Yu 0002, Dian Yu 0001, Baolin Peng, Hongwei Wang 0010, Dan Roth 0001, Dong Yu 0001
NAACL-HLT4
2024 Deceptive Semantic Shortcuts on Reasoning Chains: How Far Can Models Go without Hallucination?
abstract
Bangzheng Li, Ben Zhou, Fei Wang, Xingyu Fu, Dan Roth, Muhao Chen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Bangzheng Li, Ben Zhou, Fei Wang 0060, Dan Roth 0001, Muhao Chen 0001
NAACL-HLT2
2023 Generic Temporal Reasoning with Differential Analysis and Explanation
abstract
Temporal reasoning is the task of predicting temporal relations of event pairs.While temporal reasoning models can perform reasonably well on in-domain benchmarks, we have little idea of these systems' generalizability due to existing datasets' limitations.In this work, we introduce a novel task named TODAY that bridges this gap with temporal differential analysis, which as the name suggests, evaluates whether systems can correctly understand the effect of incremental changes.Specifically, TODAY introduces slight contextual changes for given event pairs, and systems are asked to tell how this subtle contextual change would affect relevant temporal relation distributions.To facilitate learning, TODAY also annotates human explanations.We show that existing models, including GPT-3.5, drop to random guessing on TODAY, suggesting that they heavily rely on spurious information rather than proper reasoning for temporal predictions.On the other hand, we show that TODAY's supervision style and explanation annotations can be used in joint learning, encouraging models to use more appropriate signals during training and thus outperform across several benchmarks.TODAY can also be used to train models to solicit incidental supervision from noisy sources such as GPT-3.5, thus moving us more toward the goal of generic temporal reasoning systems.
Yu Feng 0013, Ben Zhou, Haoyu Wang 0005, Helen Jin, Dan Roth 0001
ACL (1)2
2022 There's a Time and Place for Reasoning Beyond the Image
abstract
Images are often more significant than only the pixels to human eyes, as we can infer, associate, and reason with contextual information from other sources to establish a more complete picture. For example, in Figure This reasoning could provide the time and place the image was taken, which will help us in subsequent tasks, such as automatic storyline construction, correction of image source in intended effect photographs, and upper-stream processing such as image clustering for certain location or time.
Ben Zhou, Ishaan Preetam Chandratreya, Carl Vondrick, Dan Roth 0001
ACL (1)2
2022 A Meta-framework for Spatiotemporal Quantity Extraction from Text
abstract
News events are often associated with quantities (e.g., the number of COVID-19 patients or the number of arrests in a protest), and it is often important to extract their type, time, and location from unstructured text in order to analyze these quantity events.This paper thus formulates the NLP problem of spatiotemporal quantity extraction, and proposes the first meta-framework for solving it.This meta-framework contains a formalism that decomposes the problem into several information extraction tasks, a shareable crowdsourcing pipeline, and transformer-based baseline models.We demonstrate the meta-framework in three domains-the COVID-19 pandemic, Black Lives Matter protests, and 2020 California wildfires-to show that the formalism is general and extensible, the crowdsourcing pipeline facilitates fast and high-quality data annotation, and the baseline system can handle spatiotemporal quantity extraction well enough to be practically useful.We release all resources for future research on this topic.1
Qiang Ning, Ben Zhou, Hao Wu 0034, Haoruo Peng, Chuchu Fan, Matt Gardner 0001
ACL (1)2
2022 Learning to Decompose: Hypothetical Question Decomposition Based on Comparable Texts
abstract
Explicit decomposition modeling, which involves breaking down complex tasks into more straightforward and often more interpretable sub-tasks, has long been a central theme in developing robust and interpretable NLU systems.However, despite the many datasets and resources built as part of this effort, the majority have small-scale annotations and limited scope, which is insufficient to solve general decomposition tasks.In this paper, we look at large-scale intermediate pre-training of decomposition-based transformers using distant supervision from comparable texts, particularly large-scale parallel news.We show that with such intermediate pre-training, developing robust decomposition-based models for a diverse range of tasks becomes more feasible.For example, on semantic parsing, our model, DECOMPT5, improves 20% to 30% on two datasets, Overnight and TORQUE, over the baseline language model.We further use DECOMPT5 to build a novel decompositionbased QA system named DECOMPENTAIL, improving over state-of-the-art models, including GPT-3, on both HotpotQA and StrategyQA by 8% and 4%, respectively.
Ben Zhou, Kyle Richardson 0001, Xiaodong Yu 0003, Dan Roth 0001
EMNLP1
2022 End-to-End Chinese Speaker Identification
abstract
Speaker identification (SI) in texts aims to identify the speaker(s) for each utterance in texts.Previous studies divide SI into several sub-tasks (e.g., quote extraction, named entity recognition, gender identification, and coreference resolution).However, we are still far from solving these sub-tasks, making SI systems that rely on them seriously suffer from error propagation.End-to-end SI systems, on the other hand, are not limited by individual modules, but suffer from insufficient training data from the existing small-scale datasets.To make large end-to-end models possible, we design a new annotation guideline that regards SI as span extraction from the local context, and we annotate by far the largest SI dataset for Chinese named CSI based on eighteen novels.Viewing SI as a span extraction task also introduces the possibility of applying existing storng extractive machine reading comprehension (MRC) baselines.Surprisingly, simply using such a baseline without human-annotated character names and carefully designed rules, we can already achieve performance comparable or better than those of previous state-of-the-art SI methods on all public SI datasets for Chinese.Furthermore, we show that our dataset can serve as additional training data for existing benchmarks, which leads to further gains (up to 6.5% in accuracy).Finally, using CSI as a clean source, we design an effective self-training paradigm to continuously leverage hundreds of unlabeled novels.
Dian Yu 0001, Ben Zhou, Dong Yu 0001
NAACL-HLT2
2021 Cross-lingual Entity Alignment with Incidental Supervision
abstract
Much research effort has been put to multilingual knowledge graph (KG) embedding methods to address the entity alignment task, which seeks to match entities in different languagespecific KGs that refer to the same real-world object.Such methods are often hindered by the insufficiency of seed alignment provided between KGs.Therefore, we propose an incidentally supervised model, JEANS , which jointly represents multilingual KGs and text corpora in a shared embedding scheme, and seeks to improve entity alignment with incidental supervision signals from text.JEANS first deploys an entity grounding process to combine each KG with the monolingual text corpus.Then, two learning processes are conducted: (i) an embedding learning process to encode the KG and text of each language in one embedding space, and (ii) a selflearning based alignment learning process to iteratively induce the matching of entities and that of lexemes between embeddings.Experiments on benchmark datasets show that JEANS leads to promising improvement on entity alignment with incidental supervision, and significantly outperforms state-of-the-art methods that solely rely on internal information of KGs. 1 * Indicating equal contributions.
Muhao Chen 0001, Ben Zhou, Dan Roth 0001
EACL3
2021 Temporal Reasoning on Implicit Events from Distant Supervision
abstract
Ben Zhou, Kyle Richardson, Qiang Ning, Tushar Khot, Ashish Sabharwal, Dan Roth. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Ben Zhou, Kyle Richardson 0001, Qiang Ning, Tushar Khot, Ashish Sabharwal, Dan Roth 0001
NAACL-HLT1
2020 Temporal Common Sense Acquisition with Minimal Supervision
abstract
Temporal common sense (e.g., duration and frequency of events) is crucial for understanding natural language.However, its acquisition is challenging, partly because such information is often not expressed explicitly in text, and human annotation on such concepts is costly.This work proposes a novel sequence modeling approach that exploits explicit and implicit mentions of temporal common sense, extracted from a large corpus, to build TACOLM, 1 a temporal common sense language model.Our method is shown to give quality predictions of various dimensions of temporal common sense (on UDST and a newly collected dataset from Real-News).It also produces representations of events for relevant tasks such as duration comparison, parent-child relations, event coreference and temporal QA (on TimeBank, HiEVE and MCTACO) that are better than using the standard BERT.Thus, it will be an important component of temporal NLP.
Ben Zhou, Qiang Ning, Daniel Khashabi, Dan Roth 0001
ACL1
2019 "Going on a vacation" takes longer than "Going for a walk": A Study of Temporal Commonsense Understanding
abstract
Ben Zhou, Daniel Khashabi, Qiang Ning, Dan Roth. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Ben Zhou, Daniel Khashabi, Qiang Ning, Dan Roth 0001
EMNLP/IJCNLP (1)1
2018 Zero-Shot Open Entity Typing as Type-Compatible Grounding
abstract
The problem of entity-typing has been studied predominantly in supervised learning fashion, mostly with task-specific annotations (for coarse types) and sometimes with distant supervision (for fine types).While such approaches have strong performance within datasets, they often lack the flexibility to transfer across text genres and to generalize to new type taxonomies.In this work we propose a zero-shot entity typing approach that requires no annotated data and can flexibly identify newly defined types.Given a type taxonomy defined as Boolean functions of FREEBASE "types", we ground a given mention to a set of type-compatible Wikipedia entries and then infer the target mention's types using an inference algorithm that makes use of the types of these entries.We evaluate our system on a broad range of datasets, including standard fine-grained and coarse-grained entity typing datasets, and also a dataset in the biological domain.Our system is shown to be competitive with state-of-theart supervised NER systems and outperforms them on out-of-domain datasets.We also show that our system significantly outperforms other zero-shot fine typing systems.
Ben Zhou, Daniel Khashabi, Chen-Tse Tsai, Dan Roth 0001
EMNLP1
2018 CogCompNLP: Your Swiss Army Knife for NLP
Daniel Khashabi, Mark Sammons, Ben Zhou, Tom Redman, Christos Christodoulopoulos 0001, Vivek Srikumar, Nick Rizzolo, Lev-Arie Ratinov, Guanheng Luo, Quang Do, Chen-Tse Tsai, Subhro Roy, Stephen Mayhew 0001, Zhili Feng, John Wieting, Xiaodong Yu 0003, Yangqiu Song, Shashank Gupta 0007, Shyam Upadhyay, Naveen Arivazhagan, Qiang Ning, Shaoshi Ling, Dan Roth 0001
LREC3