EDBT 2026 Demo / reviewers in the wild / expert
Zhenting Qi
dblp:329/2118
· DBLP profile ↗
15ranked-venue papers
5as first author
15since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 5 first-author · 15 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Generalizing Trust: Weak-to-Strong Trustworthiness in Language ModelsabstractAs large language models continue to advance, ensuring their trustworthiness is critical.However, inaccessible real-world ground truth labels pose a significant challenge in high-stakes domains.Recent studies have highlighted weak-to-strong generalization, where a strong model trained only on a weak model's labels surpasses the weak model in task performance.Yet, whether critical trustworthiness properties such as robustness, fairness, and privacy can generalize similarly remains an open question.This is the first work to study this question by examining if a stronger model can enhance trustworthiness when fine-tuned on a weaker model's labels, a paradigm we term weak-tostrong trustworthiness.To address this, we introduce two fundamental fine-tuning strategies that leverage trustworthiness regularization during the fine-tuning of the weak model and the weak-to-strong transfer.Our experimental evaluation on real-world datasets reveals that while some trustworthiness properties, such as fairness, adversarial robustness, and OOD robustness, show significant improvement in trustworthiness generalization when both models were regularized, others, like privacy, do not exhibit signs of weak-to-strong trustworthiness.Our results highlight the potential of weak-tostrong trustworthiness as a practical pathway for enhancing the trustworthiness of increasingly capable AI systems, even under imperfect real-world conditions. Lillian Sun, Martin Pawelczyk, Zhenting Qi, Aounon Kumar, Himabindu Lakkaraju |
ACL (1) | 3 |
| 2025 | MuTIS: Enhancing Reasoning Efficiency through Multi Turn Intervention Sampling in Reinforcement LearningabstractRecently, large reasoning models (LRMs) have demonstrated state-of-the-art performance across a wide range of benchmarks.However, a common challenge for these models is the "overthinking" problem, which leads to excessive reasoning steps and significant computational overhead.Furthermore, the issues with long Chain-of-Thought (CoT) are especially pronounced in smaller models (≤ 3B parameters).Aside from producing excessively verbose "reflection words", they often exhibit repetition and get trapped in unproductive generation loops.Existing solutions typically involve either using flexible reasoning chains as training data or leveraging the model's latent space to bypass intermediate reasoning steps, but none of these methods have considered directly optimizing reasoning trajectories during the sampling phase of training.In our work, we introduce the Multi-Turn Intervention Sampling Framework (MuTIS).Our framework leverages multi-turn interventions to produce concise reasoning chains.It fine-tunes reasoning models through reinforcement learning, demonstrably breaking the accuracy-efficiency trade-off.It also demonstrates strong scalability, exhibiting excellent performance on 7B models. Wenshuo Zhao, Haoxing Zhai, Xinyu Qiu, Zhenting Qi, Shuhe Li, Linchao Zhu |
EMNLP | 4 |
| 2025 | Quantifying Generalization Complexity for Large Language ModelsabstractWhile large language models (LLMs) have shown exceptional capabilities in understanding complex queries
and performing sophisticated tasks, their generalization abilities are often deeply entangled with
memorization, necessitating more precise evaluation.
To address this challenge, we introduce Scylla, a dynamic evaluation framework that quantitatively measures the generalization abilities of LLMs. Scylla disentangles generalization from memorization via assessing model performance on both in-distribution (ID) and out-of-distribution (OOD) data through 20 tasks across 5 levels of complexity.
Through extensive experiments, we uncover a non-monotonic relationship between task complexity and the performance gap between ID and OOD
data, which we term the generalization valley.
Specifically, this phenomenon reveals a critical threshold---referred to
as critical complexity---where reliance on non-generalizable behavior peaks, indicating the
upper bound of LLMs' generalization capabilities.
As model size increases, the critical complexity shifts toward higher levels of task complexity,
suggesting that larger models can handle more complex reasoning tasks before over-relying on
memorization.
Leveraging Scylla and the concept of critical complexity, we benchmark 28 LLMs including
both open-sourced models such as LLaMA and Qwen families, and closed-sourced models like Claude and
GPT, providing a more robust evaluation and establishing a clearer
understanding of LLMs' generalization capabilities. Zhenting Qi, Hongyin Luo, Xuliang Huang, Zhuokai Zhao, Yibo Jiang, Xiangjun Fan, Himabindu Lakkaraju, James R. Glass |
ICLR | 1 |
| 2025 | Mutual Reasoning Makes Smaller LLMs Stronger Problem-SolverabstractThis paper introduces rStar, a self-play mutual reasoning approach that significantly improves reasoning capabilities of small language models (SLMs) without fine-tuning or superior models. rStar decouples reasoning into a self-play mutual generation-discrimination process. First, a target SLM augments the Monte Carlo Tree Search (MCTS) with a rich set of human-like reasoning actions to construct higher quality reasoning trajectories. Next, another SLM, with capabilities similar to the target SLM, acts as a discriminator to verify each trajectory generated by the target SLM. The mutually agreed reasoning trajectories are considered mutual consistent, thus are more likely to be correct. Extensive experiments across five SLMs demonstrate rStar can effectively solve diverse reasoning problems, including GSM8K, GSM-Hard, MATH, SVAMP, and StrategyQA. Remarkably, rStar boosts GSM8K accuracy from 12.51\% to 63.91\% for LLaMA2-7B, from 36.46\% to 81.88\% for Mistral-7B, from 74.53\% to 91.13\% for LLaMA3-8B-Instruct. Code is available at https://github.com/zhentingqi/rStar. Zhenting Qi, Mingyuan Ma, Jiahang Xu, Li Lyna Zhang, Fan Yang 0024, Mao Yang 0004 |
ICLR | 1 |
| 2025 | Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation SystemsabstractRetrieval-Augmented Generation (RAG) improves pre-trained models by incorporating external knowledge at test time to enable customized adaptation.
We study the risk of datastore leakage in Retrieval-In-Context RAG Language Models (LMs). We show that an adversary can exploit LMs' instruction-following capabilities to easily extract text data verbatim from the datastore of RAG systems built with instruction-tuned LMs via prompt injection.
The vulnerability exists for a wide range of modern LMs that span Llama2, Mistral/Mixtral, Vicuna, SOLAR, WizardLM, Qwen1.5, and Platypus2, and the exploitability exacerbates as the model size scales up.
We also study multiple effects of RAG setup on the extractability of data, indicating that following unexpected instructions to regurgitate data can be an outcome of failure in effectively utilizing contexts for modern LMs, and further show that such vulnerability can be greatly mitigated by position bias elimination strategies.
Extending our study to production RAG models, GPTs, we design an attack that can cause datastore leakage with a near-perfect success rate on 25 randomly selected customized GPTs with at most 2 queries, and we extract text data verbatim at a rate of 41\% from a book of 77,000 words and 3\% from a corpus of 1,569,000 words by prompting the GPTs with only 100 queries generated by themselves. Zhenting Qi, Hanlin Zhang 0002, Eric P. Xing, Sham M. Kakade, Himabindu Lakkaraju |
ICLR | 1 |
| 2025 | Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive SearchabstractLarge language models (LLMs) have demonstrated remarkable reasoning capabilities across diverse domains. Recent studies have shown that increasing test-time computation enhances LLMs' reasoning capabilities. This typically involves extensive sampling at inference time guided by an external LLM verifier, resulting in a two-player system. Despite external guidance, the effectiveness of this system demonstrates the potential of a single LLM to tackle complex tasks. Thus, we pose a new research problem: *Can we internalize the searching capabilities to fundamentally enhance the reasoning abilities of a single LLM?* This work explores an orthogonal direction focusing on post-training LLMs for autoregressive searching (*i.e.,* an extended reasoning process with self-reflection and self-exploration of new strategies). To achieve this, we propose the Chain-of-Action-Thought (COAT) reasoning and a two-stage training paradigm: 1) a small-scale format tuning stage to internalize the COAT reasoning format and 2) a large-scale self-improvement stage leveraging reinforcement learning. Our approach results in Satori, a 7B LLM trained on open-source models and data. Extensive empirical evaluations demonstrate that Satori achieves state-of-the-art performance on mathematical reasoning benchmarks while exhibits strong generalization to out-of-domain tasks. Code, data, and models are fully open-sourced. Maohao Shen, Guangtao Zeng, Zhenting Qi, Zhang-Wei Hong, Zhenfang Chen, Gregory W. Wornell, Subhro Das, David D. Cox, Chuang Gan 0001 |
ICML | 3 |
| 2025 | EvoLM: In Search of Lost Training Dynamics for Language Model ReasoningabstractModern language model (LM) training has been divided into multiple stages, making it difficult for downstream developers to evaluate the impact of design choices made at each stage.
We present EvoLM, a model suite that enables systematic and transparent analysis of LMs' training dynamics across pre-training, continued pre-training, supervised fine-tuning, and reinforcement learning.
By training over 100 LMs with 1B and 4B parameters from scratch, we rigorously evaluate both upstream (language modeling) and downstream (problem-solving) reasoning capabilities, including considerations of both in-domain and out-of-domain generalization.
Key insights highlight the diminishing returns from excessive pre-training and post-training, the importance and practices of mitigating forgetting during domain-specific continued pre-training, the crucial role of continued pre-training in bridging pre-training and post-training phases, and various intricate trade-offs when configuring supervised fine-tuning and reinforcement learning.
To facilitate open research and reproducibility, we release all pre-trained and post-trained models, training datasets for all stages, and our entire training and evaluation pipeline. Zhenting Qi, Fan Nie, Alexandre Alahi, James Zou 0001, Himabindu Lakkaraju, Yilun Du, Eric P. Xing, Sham M. Kakade, Hanlin Zhang 0002 |
NeurIPS | 1 |
| 2025 | Measuring the Faithfulness of Thinking Drafts in Large Reasoning ModelsabstractLarge Reasoning Models (LRMs) have significantly enhanced their capabilities in complex problem-solving by introducing a thinking draft that enables multi-path Chain-of-Thought explorations before producing final answers.
Ensuring the faithfulness of these intermediate reasoning processes is crucial for reliable monitoring, interpretation, and effective control. In this paper, we propose a systematic counterfactual intervention framework to rigorously evaluate *thinking draft faithfulness*.
Our approach focuses on two complementary dimensions:
**(1) Intra-Draft Faithfulness**, which assesses whether individual reasoning steps causally influence subsequent steps and the final draft conclusion through counterfactual step insertions; and
**(2) Draft-to-Answer Faithfulness**, which evaluates whether final answers are logically consistent with and dependent on the thinking draft, by perturbing the draft’s concluding logic.
We conduct extensive experiments across six state-of-the-art LRMs.
Our findings show that current LRMs demonstrate selective faithfulness to intermediate reasoning steps and frequently fail to faithfully align with the draft conclusions.
These results underscore the need for more faithful and interpretable reasoning in advanced LRMs. Zidi Xiong, Zhenting Qi, Himabindu Lakkaraju |
NeurIPS | 3 |
| 2024 | FOLIO: Natural Language Reasoning with First-Order LogicabstractSimeng Han, Hailey Schoelkopf, Yilun Zhao, Zhenting Qi, Martin Riddell, Wenfei Zhou, James Coady, David Peng, Yujie Qiao, Luke Benson, Lucy Sun, Alexander Wardle-Solano, Hannah Szabó, Ekaterina Zubova, Matthew Burtell, Jonathan Fan, Yixin Liu, Brian Wong, Malcolm Sailor, Ansong Ni, Linyong Nan, Jungo Kasai, Tao Yu, Rui Zhang, Alexander Fabbri, Wojciech Maciej Kryscinski, Semih Yavuz, Ye Liu, Xi Victoria Lin, Shafiq Joty, Yingbo Zhou, Caiming Xiong, Rex Ying, Arman Cohan, Dragomir Radev. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Simeng Han, Hailey Schoelkopf, Yilun Zhao 0001, Zhenting Qi, Martin Riddell, Wenfei Zhou, James Coady, David Peng, Yujie Qiao, Luke Benson, Lucy Sun, Alexander Wardle-Solano, Hannah Szabó, Ekaterina Zubova, Matthew Burtell, Jonathan Fan 0001, Yixin Liu 0003, Malcolm Sailor, Ansong Ni, Linyong Nan, Jungo Kasai, Tao Yu 0009, Rui Zhang 0037, Alexander R. Fabbri, Wojciech Kryscinski, Semih Yavuz, Ye Liu 0006, Xi Victoria Lin, Shafiq R. Joty, Yingbo Zhou 0002, Caiming Xiong, Rex Ying, Arman Cohan, Dragomir R. Radev |
EMNLP | 4 |
| 2024 | Constrained Human-AI Cooperation: An Inclusive Embodied Social Intelligence ChallengeabstractWe introduce Constrained Human-AI Cooperation (CHAIC), an inclusive embodied social intelligence challenge designed to test social perception and cooperation in embodied agents. In CHAIC, the goal is for an embodied agent equipped with egocentric observations to assist a human who may be operating under physical constraints—e.g., unable to reach high places or confined to a wheelchair—in performing common household or outdoor tasks as efficiently as possible. To achieve this, a successful helper must: (1) infer the human's intents and constraints by following the human and observing their behaviors (social perception), and (2) make a cooperative plan tailored to the human partner to solve the task as quickly as possible, working together as a team (cooperative planning). To benchmark this challenge, we create four new agents with real physical constraints and eight long-horizon tasks featuring both indoor and outdoor scenes with various constraints, emergency events, and potential risks. We benchmark planning- and learning-based baselines on the challenge and introduce a new method that leverages large language models and behavior modeling. Empirical evaluations demonstrate the effectiveness of our benchmark in enabling systematic assessment of key aspects of machine social intelligence. Our benchmark and code are publicly available at https://github.com/UMass-Foundation-Model/CHAIC. Weihua Du, Qiushi Lyu, Jiaming Shan, Zhenting Qi, Sunli Chen, Andi Peng, Tianmin Shu, Kwonjoon Lee, Behzad Dariush, Chuang Gan 0001 |
NeurIPS | 4 |
| 2023 | RobuT: A Systematic Study of Table QA Robustness Against Human-Annotated Adversarial PerturbationsabstractYilun Zhao, Chen Zhao, Linyong Nan, Zhenting Qi, Wenlin Zhang, Xiangru Tang, Boyu Mi, Dragomir Radev. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yilun Zhao 0001, Chen Zhao 0013, Linyong Nan, Zhenting Qi, Xiangru Tang, Boyu Mi, Dragomir R. Radev |
ACL (1) | 4 |
| 2023 | LoFT: Enhancing Faithfulness and Diversity for Table-to-Text Generation via Logic Form ControlabstractLogical Table -to-Text (LT2T) generation is tasked with generating logically faithful sentences from tables.There currently exists two challenges in the field: 1) Faithfulness: how to generate sentences that are factually correct given the table content; 2) Diversity: how to generate multiple sentences that offer different perspectives on the table.This work proposes LOFT, which utilizes logic forms as fact verifiers and content planners to control LT2T generation.Experimental results on the LOGICNLG dataset demonstrate that LOFT is the first model that addresses unfaithfulness and lack of diversity issues simultaneously.Our code is publicly available at https: //github.com/Yale-LILY/LoFT. Yilun Zhao 0001, Zhenting Qi, Linyong Nan, Lorenzo Jaime Yu Flores, Dragomir R. Radev |
EACL | 2 |
| 2023 | QTSumm: Query-Focused Summarization over Tabular DataabstractYilun Zhao, Zhenting Qi, Linyong Nan, Boyu Mi, Yixin Liu, Weijin Zou, Simeng Han, Ruizhe Chen, Xiangru Tang, Yumo Xu, Dragomir Radev, Arman Cohan. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Yilun Zhao 0001, Zhenting Qi, Linyong Nan, Boyu Mi, Yixin Liu 0003, Weijin Zou, Simeng Han, Ruizhe Chen, Xiangru Tang, Yumo Xu, Dragomir R. Radev, Arman Cohan |
EMNLP | 2 |
| 2022 | ReasTAP: Injecting Table Reasoning Skills During Pre-training via Synthetic Reasoning ExamplesabstractReasoning over tabular data requires both table structure understanding and a broad set of table reasoning skills.Current models with tablespecific architectures and pre-training methods perform well on understanding table structures, but they still struggle with tasks that require various table reasoning skills.In this work, we develop REASTAP to show that high-level table reasoning skills can be injected into models during pre-training without a complex tablespecific architecture design.We define 7 table reasoning skills, such as numerical operation, temporal comparison, and conjunction.Each reasoning skill is associated with one example generator, which synthesizes questions over semi-structured tables according to the sampled templates.We model the table pre-training task as a sequence generation task and pretrain REASTAP to generate precise answers to the synthetic examples.REASTAP is evaluated on four benchmarks covering three downstream tasks including: 1) WIKISQL-WEAK and WIKITQ for Table Question Answering; 2) TABFACT for Table Fact Verification; and 3) LOGICNLG for Faithful Table-to-Text Generation.Experimental results demonstrate that REASTAP achieves new state-of-the-art performance on all benchmarks and delivers a significant improvement on low-resource setting. Yilun Zhao 0001, Linyong Nan, Zhenting Qi, Rui Zhang 0037, Dragomir R. Radev |
EMNLP | 3 |
| 2022 | Weakly Supervised Two-Stage Training Scheme for Deep Video Fight Detection ModelabstractFight detection in videos is an emerging deep learning application with today's prevalence of surveillance systems and streaming media. Previous work has largely relied on action recognition techniques to tackle this problem. In this paper, we propose a simple but effective method that solves the task from a new perspective: we design the fight detection model as a composition of an action-aware feature extractor and an anomaly score generator. Also, considering that collecting frame-level labels for videos is too laborious, we design a weakly supervised two-stage training scheme, where we utilize multiple-instance-learning loss calculated on video-level labels to train the score generator, and adopt the self-training technique to further improve its performance. Extensive experiments on a publicly available large-scale dataset, UBI-Fights, demonstrate the effectiveness of our method, and the performance on the dataset exceeds several previous state-of-the-art approaches. Furthermore, we collect a new dataset, VFD-2000, that specializes in video fight detection, with a larger scale and more scenarios than existing datasets. The implementation of our method and the proposed dataset is available at https://github.com/Hepta-Col/VideoFightDetection. Zhenting Qi, Ruike Zhu, Zheyu Fu, Wenhao Chai, Volodymyr V. Kindratenko |
ICTAI | 1 |