EDBT 2026 Demo / reviewers in the wild / expert
Ruiqi Zhong
dblp:222/3024
· DBLP profile ↗
19ranked-venue papers
7as first author
14since 2021 · last 2025
0000-0002-3153-4027ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 7 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Language Models Learn to Mislead Humans via RLHFabstractLanguage models (LMs) can produce errors that are hard to detect for humans, especially when the task is complex.
RLHF, the most popular post-training method, may exacerbate this problem: to achieve higher rewards, LMs might get better at convincing humans that they are right even when they are wrong. We study this phenomenon under a standard RLHF pipeline, calling it ``U-Sophistry'' since it is \textbf{U}nintended by model developers. Specifically, we ask time-constrained (e.g., 3-10 minutes) human subjects to evaluate the correctness of model outputs and calculate humans' accuracy against gold labels. On a question-answering task (QuALITY) and programming task (APPS), RLHF makes LMs better at convincing our subjects but not at completing the task correctly. RLHF also makes the model harder to evaluate: our subjects' false positive rate increases by 24.1% on QuALITY and 18.3% on APPS.
Finally, we show that probing, a state-of-the-art approach for detecting \textbf{I}ntended Sophistry (e.g.~backdoored LMs), does not generalize to U-Sophistry. Our results highlight an important failure mode of RLHF and call for more research in assisting humans to align them. Jiaxin Wen, Ruiqi Zhong, Akbir Khan, Ethan Perez, Jacob Steinhardt, Minlie Huang, Samuel R. Bowman, He He 0001, Shi Feng 0005 |
ICLR | 2 |
| 2025 | Language Models can Categorize System Inputs for Performance AnalysisabstractDominic Sobhani, Ruiqi Zhong, Edison Marrese-Taylor, Keisuke Sakaguchi, Yutaka Matsuo. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Dominic Sobhani, Ruiqi Zhong, Edison Marrese-Taylor, Keisuke Sakaguchi, Yutaka Matsuo |
NAACL (Long Papers) | 2 |
| 2024 | Learning Task Decomposition to Assist Humans in Competitive ProgrammingabstractWhen using language models (LMs) to solve complex problems, humans might struggle to understand the LM-generated solutions and repair the flawed ones.To assist humans in repairing them, we propose to automatically decompose complex solutions into multiple simpler pieces that correspond to specific subtasks.We introduce a novel objective for learning task decomposition, termed assistive value (AssistV), which measures the feasibility and speed for humans to repair the decomposed solution.We collect a dataset of human repair experiences on different decomposed solutions.Utilizing the collected data as in-context examples, we then learn to critique, refine, and rank decomposed solutions to improve AssistV.We validate our method under competitive programming problems: under 177 hours of human study, our method enables non-experts to solve 33.3% more problems, speeds them up by 3.3x, and empowers them to match unassisted experts. Jiaxin Wen, Ruiqi Zhong, Pei Ke, Zhihong Shao, Hongning Wang, Minlie Huang |
ACL (1) | 2 |
| 2024 | Describing Differences in Image Sets with Natural LanguageabstractHow do two sets of images differ? Discerning set-level differences is crucial for understanding model behaviors and analyzing datasets, yet manually sifting through thousands of images is impractical. To aid in this discovery process, we explore the task of automatically describing the differences between two sets of images, which we term Set Difference Captioning. This task takes in image sets$\mathcal{D}_{A}$and$\mathcal{D}_{B}$, and outputs a description that is more often true on$\mathcal{D}_{A}$than$\mathcal{D}_{B}$. We outline a two-stage approach that first proposes candidate difference descriptions from image sets and then re-ranks the candidates by checking how well they can differentiate the two sets. We introduce VisDiff, which first captions the images and prompts a language model to propose candidate descriptions, then re-ranks these descriptions using CLIP. To evaluate VisDiff, we collect VisDiffBench, a dataset with 187 paired image sets with ground truth difference descriptions. We apply VisDiff to various domains, such as comparing datasets (e.g., ImageNet vs. ImageNetV2), comparing classification models (e.g., zero-shot CLIP vs. supervised ResNet), characterizing differences between generative models (e.g., StableDiffusionV1 and V2), and discovering what makes images memorable. Using VisDiff, we are able to find interesting and previously unknown differences in datasets and models, demonstrating its utility in revealing nuanced insights.11Project page available at https:/understanding-visual-datasets.github.io/VisDiff-website/. Lisa Dunlap, Ruiqi Zhong, Trevor Darrell, Jacob Steinhardt, Joseph Gonzalez 0001, Serena Yeung-Levy |
CVPR | 4 |
| 2024 | Do Models Explain Themselves? Counterfactual Simulatability of Natural Language ExplanationsabstractLarge language models (LLMs) are trained to imitate humans to explain human decisions. However, do LLMs explain themselves? Can they help humans build mental models of how LLMs process different inputs? To answer these questions, we propose to evaluate $\textbf{counterfactual simulatability}$ of natural language explanations: whether an explanation can enable humans to precisely infer the model’s outputs on diverse counterfactuals of the explained input. For example, if a model answers ”$\textit{yes}$” to the input question ”$\textit{Can eagles fly?}$” with the explanation ”$\textit{all birds can fly}$”, then humans would infer from the explanation that it would also answer ”$\textit{yes}$” to the counterfactual input ”$\textit{Can penguins fly?}$”. If the explanation is precise, then the model’s answer should match humans’ expectations. We implemented two metrics based on counterfactual simulatability: precision and generality. We generated diverse counterfactuals automatically using LLMs. We then used these metrics to evaluate state-of-the-art LLMs (e.g., GPT-4) on two tasks: multi-hop factual reasoning and reward modeling. We found that LLM’s explanations have low precision and that precision does not correlate with plausibility. Therefore, naively optimizing human approvals (e.g., RLHF) may be insufficient. Yanda Chen, Ruiqi Zhong, Narutatsu Ri, Chen Zhao 0013, He He 0001, Jacob Steinhardt, Kathy McKeown |
ICML | 2 |
| 2024 | Explaining Datasets in Words: Statistical Models with Natural Language ParametersabstractTo make sense of massive data, we often first fit simplified models and then interpret the parameters; for example, we cluster the text embeddings and then interpret the mean parameters of each cluster.
However, these parameters are often high-dimensional and hard to interpret.
To make model parameters directly interpretable, we introduce a family of statistical models---including clustering, time series, and classification models---parameterized by *natural language predicates*.
For example, a cluster of text about COVID could be parameterized by the predicate ``*discusses COVID*''.
To learn these statistical models effectively, we develop a model-agnostic algorithm that optimizes continuous relaxations of predicate parameters with gradient descent and discretizes them by prompting language models (LMs).
Finally, we apply our framework to a wide range of problems: taxonomizing user chat dialogues, characterizing how they evolve across time, finding categories where one language model is better than the other, clustering math problems based on subareas, and explaining visual features in memorable images.
Our framework is highly versatile, applicable to both textual and visual domains, can be easily steered to focus on specific properties (e.g. subareas), and explains sophisticated concepts that classical methods (e.g. n-gram analysis) struggle to produce. Ruiqi Zhong, Daniel Klein 0001, Jacob Steinhardt |
NeurIPS | 1 |
| 2023 | Goal-Driven Explainable Clustering via Language DescriptionsabstractUnsupervised clustering is widely used to explore large corpora, but existing formulations neither consider the users' goals nor explain clusters' meanings.We propose a new task formulation, "Goal-Driven Clustering with Explanations" (GOALEX), which represents both the goal and the explanations as free-form language descriptions.For example, to categorize the errors made by a summarization system, the input to GOALEX is a corpus of annotatorwritten comments for system-generated summaries and a goal "cluster the comments based on why the annotators think the summary is imperfect.";the outputs are text clusters each with an explanation ("this cluster mentions that the summary misses important context information."),which relates to the goal and accurately explains which comments should (not) belong to a cluster.To tackle GOALEX, we prompt a language model with "[corpus subset] + [goal] + Brainstorm a list of explanations each representing a cluster.";then we classify whether each sample belongs to a cluster based on its explanation; finally, we use integer linear programming to select a subset of candidate clusters to cover most samples while minimizing overlaps.Under both automatic and human evaluation on corpora with or without labels, our method produces more accurate and goalrelated explanations than prior methods. Zihan Wang 0001, Jingbo Shang, Ruiqi Zhong |
EMNLP | 3 |
| 2023 | Non-Programmers Can Label Programs Indirectly via Active Examples: A Case Study with Text-to-SQLabstractCan non-programmers annotate natural language utterances with complex programs that represent their meaning?We introduce APEL, a framework in which non-programmers select among candidate programs generated by a seed semantic parser (e.g., Codex).Since they cannot understand the candidate programs, we ask them to select indirectly by examining the programs' input-ouput examples.For each utterance, APEL actively searches for a simple input on which the candidate programs tend to produce different outputs.It then asks the nonprogrammers only to choose the appropriate output, thus allowing us to infer which program is correct and could be used to fine-tune the parser.As a case study, we recruited human non-programmers to use APEL to re-annotate SPIDER, a text-to-SQL dataset.Our approach achieved the same annotation accuracy as the original expert annotators (75%) and exposed many subtle errors in the original annotations.Utterance u: Find the first name of students who have both cat and dog pets.SELECT fname FROM Student WHERE StuID IN (SELECT T1.stuid FROM student AS T1 JOIN has_pet AS T2 ON T1.stuid = T2.stuidJOIN pets AS T3 ON T3.petid = T2.petidWHERE T3.pettype = 'cat' INTERSECT SELECT T1.stuid FROM student AS T1 JOIN has_pet AS T2 ON T1.stuid = T2.stuidJOIN pets AS T3 ON T3.petid = T2.petidWHERE T3.pettype = 'dog') SELECT t1.fname FROM student AS t1 JOIN has_pet AS t2 ON t1.stuid = t2.stuidJOIN pets AS t3 ON t3.petid = t2.petidWHERE t3.pettype = 'cat' INTERSECT SELECT t1.fname FROM student AS t1 JOIN has_pet AS t2 ON t1.stuid = t2.stuidJOIN pets AS t3 ON t3.petid = t2.petidWHERE t3.pettype = 'dog' Ruiqi Zhong, Charlie Snell, Daniel Klein 0001, Jason Eisner |
EMNLP | 1 |
| 2023 | InCoder: A Generative Model for Code Infilling and Synthesis
Daniel Fried, Armen Aghajanyan, Jessy Lin, Sida I. Wang, Eric Wallace, Freda Shi, Ruiqi Zhong, Scott Yih, Luke Zettlemoyer, Mike Lewis |
ICLR | 7 |
| 2023 | DS-1000: A Natural and Reliable Benchmark for Data Science Code GenerationabstractWe introduce DS-1000, a code generation benchmark with a thousand data science problems spanning seven Python libraries, such as Numpy and Pandas. Compared to prior works, DS-1000 incorporates three core features. First, our problems reflect diverse, realistic, and practical use cases since we collected them from StackOverflow. Second, our automatic evaluation is highly specific (reliable) – across all Codex-002-predicted solutions that our evaluation accepts, only 1.8% of them are incorrect; we achieve this with multi-criteria metrics, checking both functional correctness by running test cases and surface-form constraints by restricting API usages or keywords. Finally, we proactively defend against memorization by slightly modifying our problems to be different from the original StackOverflow source; consequently, models cannot answer them correctly by memorizing the solutions from pre-training. The current best public system (Codex-002) achieves 43.3% accuracy, leaving ample room for improvement. We release our benchmark at https://ds1000-code-gen.github.io. Yuhang Lai, Chengxi Li 0011, Ruiqi Zhong, Luke Zettlemoyer, Scott Yih, Daniel Fried, Sida I. Wang, Tao Yu 0009 |
ICML | 5 |
| 2023 | Goal Driven Discovery of Distributional Differences via Language DescriptionsabstractExploring large corpora can generate useful discoveries but is time-consuming for humans.
We formulate a new task, D5, that automatically discovers differences between two large corpora in a goal-driven way.
The task input is a problem comprising a user-specified research goal (“*comparing the side effects of drug A and drug*”) and a corpus pair (two large collections of patients' self-reported reactions after taking each drug).
The output is a goal-related description (discovery) of how these corpora differ (patients taking drug A “*mention feelings of paranoia*” more often).
We build a D5 system, and to quantitatively evaluate its performance, we 1) build a diagnostic benchmark, SynD5, to test whether it can recover known differences between two synthetic corpora, and 2) contribute a meta-dataset, OpenD5, aggregating 675 open-ended problems ranging across business, social sciences, humanities, machine learning, and health.
With both synthetic and real datasets, we confirm that language models can leverage the user-specified goals to propose more relevant candidate discoveries, and they sometimes produce discoveries previously unknown to the authors, including demographic differences in discussion topics, political stances in speech, insights in commercial reviews, and error patterns in NLP models.
Finally, we discuss the limitations of the current D5 system, which discovers correlation rather than causation and has the potential to reinforce societal biases when misused; therefore, practitioners should treat the outputs of our system with caution. Ruiqi Zhong, Peter Zhang, Steve Li, Jinwoo Ahn, Daniel Klein 0001, Jacob Steinhardt |
NeurIPS | 1 |
| 2022 | Meta-learning via Language Model In-context TuningabstractThe goal of meta-learning is to learn to adapt to a new task with only a few labeled examples.Inspired by the recent progress in large language models, we propose in-context tuning (ICT), which recasts task adaptation and prediction as a simple sequence prediction problem: to form the input sequence, we concatenate the task instruction, labeled in-context examples, and the target input to predict; to metatrain the model to learn from in-context examples, we fine-tune a pre-trained language model (LM) to predict the target label given the input sequence on a collection of tasks.We benchmark our method on two collections of text classification tasks: LAMA and Bina-ryClfs.Compared to MAML which adapts the model through gradient descent, our method leverages the inductive bias of pre-trained LMs to perform pattern matching, and outperforms MAML by an absolute 6% average AUC-ROC score on BinaryClfs, gaining more advantage with increasing model size.Compared to non-fine-tuned in-context learning (i.e.prompting a raw LM), in-context tuning meta-trains the model to learn from in-context examples.On BinaryClfs, ICT improves the average AUC-ROC score by an absolute 10%, and reduces the variance due to example ordering by 6x and example choices by 2x. Yanda Chen, Ruiqi Zhong, Sheng Zha, George Karypis, He He 0001 |
ACL (1) | 2 |
| 2022 | UnifiedSKG: Unifying and Multi-Tasking Structured Knowledge Grounding with Text-to-Text Language ModelsabstractTianbao Xie, Chen Henry Wu, Peng Shi, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong, Pengcheng Yin, Sida I. Wang, Victor Zhong, Bailin Wang, Chengzu Li, Connor Boyle, Ansong Ni, Ziyu Yao, Dragomir Radev, Caiming Xiong, Lingpeng Kong, Rui Zhang, Noah A. Smith, Luke Zettlemoyer, Tao Yu. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Tianbao Xie, Chen Henry Wu, Peng Shi 0010, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong 0005, Sida I. Wang, Victor Zhong, Bailin Wang, Chengzu Li, Connor Boyle, Ansong Ni, Ziyu Yao 0002, Dragomir R. Radev, Caiming Xiong, Lingpeng Kong, Rui Zhang 0037, Noah A. Smith, Luke Zettlemoyer, Tao Yu 0009 |
EMNLP | 4 |
| 2022 | Describing Differences between Text Distributions with Natural LanguageabstractHow do two distributions of text differ? Humans are slow at answering this, since discovering patterns might require tediously reading through hundreds of samples. We propose to automatically summarize the differences by “learning a natural language hypothesis": given two distributions $D_{0}$ and $D_{1}$, we search for a description that is more often true for $D_{1}$, e.g., “is military-related." To tackle this problem, we fine-tune GPT-3 to propose descriptions with the prompt: “[samples of $D_{0}$] + [samples of $D_{1}$] + the difference between them is \underline{\space\space\space\space}". We then re-rank the descriptions by checking how often they hold on a larger set of samples with a learned verifier. On a benchmark of 54 real-world binary classification tasks, while GPT-3 Curie (13B) only generates a description similar to human annotation 7% of the time, the performance reaches 61% with fine-tuning and re-ranking, and our best system using GPT-3 Davinci (175B) reaches 76%. We apply our system to describe distribution shifts, debug dataset shortcuts, summarize unknown tasks, and label text clusters, and present analyses based on automatically generated descriptions. Ruiqi Zhong, Charlie Snell, Daniel Klein 0001, Jacob Steinhardt |
ICML | 1 |
| 2020 | Semantic Scaffolds for Pseudocode-to-Code GenerationabstractWe propose a method for program generation based on semantic scaffolds, lightweight structures representing the high-level semantic and syntactic composition of a program.By first searching over plausible scaffolds then using these as constraints for a beam search over programs, we achieve better coverage of the search space when compared with existing techniques.We apply our hierarchical search method to the SPoC dataset for pseudocodeto-code generation, in which we are given line-level natural language pseudocode annotations and aim to produce a program satisfying execution-based test cases.By using semantic scaffolds during inference, we achieve a 10% absolute improvement in top-100 accuracy over the previous state-of-the-art.Additionally, we require only 11 candidates to reach the top-3000 performance of the previous best approach when tested against unseen problems, demonstrating a substantial improvement in efficiency. Ruiqi Zhong, Mitchell Stern, Daniel Klein 0001 |
ACL | 1 |
| 2020 | Semantic Evaluation for Text-to-SQL with Distilled Test SuitesabstractWe propose test suite accuracy to approximate semantic accuracy for Text-to-SQL models.Our method distills a small test suite of databases that achieves high code coverage for the gold query from a large number of randomly generated databases.At evaluation time, it computes the denotation accuracy of the predicted queries on the distilled test suite, hence calculating a tight upper-bound for semantic accuracy efficiently.We use our proposed method to evaluate 21 models submitted to the Spider leader board and manually verify that our method is always correct on 100 examples.In contrast, the current Spider metric leads to a 2.5% false negative rate on average and 8.1% in the worst case, indicating that test suite accuracy is needed.Our implementation, along with distilled test suites for eleven Textto-SQL datasets, is publicly available. Ruiqi Zhong, Tao Yu 0009, Daniel Klein 0001 |
EMNLP (1) | 1 |
| 2019 | Detecting and Reducing Bias in a High Stakes DomainabstractRuiqi Zhong, Yanda Chen, Desmond Patton, Charlotte Selous, Kathy McKeown. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Ruiqi Zhong, Yanda Chen, Desmond Upton Patton, Charlotte Selous, Kathy McKeown |
EMNLP/IJCNLP (1) | 1 |
| 2018 | Detecting Gang-Involved Escalation on Social Media Using ContextabstractSerina Chang, Ruiqi Zhong, Ethan Adams, Fei-Tzin Lee, Siddharth Varia, Desmond Patton, William Frey, Chris Kedzie, Kathy McKeown. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2018. Serina Chang, Ruiqi Zhong, Ethan Adams, Fei-Tzin Lee, Siddharth Varia, Desmond Upton Patton, William R. Frey, Chris Kedzie, Kathy McKeown |
EMNLP | 2 |
| 2018 | Subspace Embedding and Linear Regression with Orlicz NormabstractWe consider a generalization of the classic linear regression problem to the case when the loss is an Orlicz norm. An Orlicz norm is parameterized by a non-negative convex function G: R_+ - > R_+ with G(0) = 0: the Orlicz norm of a n-dimensional vector x is defined as |x|_G = inf{ alpha > 0 | sum_{i = 1}^n G( |x_i| / alpha ) < = 1 }. We consider the cases where the function G grows subquadratically. Our main result is based on a new oblivious embedding which embeds the column space of a given nxd matrix A with Orlicz norm into a lower dimensional space with L2 norm. Specifically, we show how to efficiently find an mxn embedding matrix S (m < n), such that for every d-dimensional vector x, we have Omega(1/(d log n)) |Ax|_G < = |SAx|_2 < = O(d^2 log n) |Ax|_G. By applying this subspace embedding technique, we show an approximation algorithm for the regression problem min_x |Ax-b|_G, up to a O( d log^2 n ) factor. As a further application of our techniques, we show how to also use them to improve on the algorithm for the Lp low rank matrix approximation problem for 1 < = p < 2. Alexandr Andoni, Chengyu Lin 0001, Ying Sheng 0004, Peilin Zhong, Ruiqi Zhong |
ICML | 5 |