Shi Feng 0005

dblp:97/1374-5 · DBLP profile ↗
← Back
21ranked-venue papers
4as first author
15since 2021 · last 2025
0009-0001-1685-8912ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 3 first-author · 15 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User Personas
abstract
Nishant Balepur, Vishakh Padmakumar, Fumeng Yang, Shi Feng, Rachel Rudinger, Jordan Lee Boyd-Graber. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Nishant Balepur, Vishakh Padmakumar, Fumeng Yang, Shi Feng 0005, Rachel Rudinger, Jordan L. Boyd-Graber
ACL (1)4
2025 A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
abstract
Nishant Balepur, Matthew Shu, Yoo Yeon Sung, Seraphina Goldfarb-Tarrant, Shi Feng, Fumeng Yang, Rachel Rudinger, Jordan Lee Boyd-Graber. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Nishant Balepur, Matthew Shu, Yoo Yeon Sung, Seraphina Goldfarb-Tarrant, Shi Feng 0005, Fumeng Yang, Rachel Rudinger, Jordan L. Boyd-Graber
EMNLP5
2025 Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats
abstract
As large language models (LLMs) grow more powerful, they also become more difficult to trust. They could be either aligned with human intentions, or exhibit "subversive misalignment" -- introducing subtle errors that bypass safety checks. Although individual errors may not immediately cause harm, each increases the risk of an eventual safety failure. With this uncertainty, model deployment often grapples with the tradeoff between ensuring safety and harnessing the capabilities of untrusted models. In this work, we introduce the ``Diffuse Risk Management'' problem, aiming to balance the average-case safety and usefulness in the deployment of untrusted models over a large sequence of tasks. We approach this problem by developing a two-level framework: the single-task level (micro-protocol) and the whole-scenario level (macro-protocol). At the single-task level, we develop various \textit{micro}-protocols that use a less capable, but extensively tested (trusted) model to harness and monitor the untrusted model. At the whole-scenario level, we find an optimal \textit{macro}-protocol that uses an adaptive estimate of the untrusted model's risk to choose between micro-protocols. To evaluate the robustness of our method, we follow \textit{control evaluations} in a code generation testbed, which involves a red team attempting to generate subtly backdoored code with an LLM whose deployment is safeguarded by a blue team. Experiment results show that our approach retains 99.6\% usefulness of the untrusted model while ensuring near-perfect safety, significantly outperforming existing deployment methods. Our approach also demonstrates robustness when the trusted and untrusted models have a large capability gap. Our findings demonstrate the promise of managing diffuse risks in the deployment of increasingly capable but untrusted LLMs.
Jiaxin Wen, Vivek Hebbar, Caleb Larson, Aryan Bhatt, Ansh Radhakrishnan, Mrinank Sharma, Henry Sleight, Shi Feng 0005, He He 0001, Ethan Perez, Buck Shlegeris, Akbir Khan
ICLR8
2025 Language Models Learn to Mislead Humans via RLHF
abstract
Language models (LMs) can produce errors that are hard to detect for humans, especially when the task is complex. RLHF, the most popular post-training method, may exacerbate this problem: to achieve higher rewards, LMs might get better at convincing humans that they are right even when they are wrong. We study this phenomenon under a standard RLHF pipeline, calling it ``U-Sophistry'' since it is \textbf{U}nintended by model developers. Specifically, we ask time-constrained (e.g., 3-10 minutes) human subjects to evaluate the correctness of model outputs and calculate humans' accuracy against gold labels. On a question-answering task (QuALITY) and programming task (APPS), RLHF makes LMs better at convincing our subjects but not at completing the task correctly. RLHF also makes the model harder to evaluate: our subjects' false positive rate increases by 24.1% on QuALITY and 18.3% on APPS. Finally, we show that probing, a state-of-the-art approach for detecting \textbf{I}ntended Sophistry (e.g.~backdoored LMs), does not generalize to U-Sophistry. Our results highlight an important failure mode of RLHF and call for more research in assisting humans to align them.
Jiaxin Wen, Ruiqi Zhong, Akbir Khan, Ethan Perez, Jacob Steinhardt, Minlie Huang, Samuel R. Bowman, He He 0001, Shi Feng 0005
ICLR9
2025 Predicting Empirical AI Research Outcomes with Language Models
abstract
Many promising-looking ideas in AI research fail to deliver, but their validation takes substantial human labor and compute. Predicting an idea's chance of success is thus crucial for accelerating empirical AI research, a skill that even expert researchers can only acquire through substantial experience. We build the first benchmark for this task and compare LMs with human experts. Concretely, given two research ideas (e.g., two jailbreaking methods), we aim to predict which will perform better on a set of benchmarks. We scrape ideas and experimental results from conference papers, yielding 1,585 human-verified idea pairs \textit{published after our base model's cut-off date} for testing, and 6,000 pairs for training. We then develop a system that combines a fine-tuned GPT-4.1 with a paper retrieval agent, and we recruit 25 human experts to compare with. In the NLP domain, our system beats human experts by a large margin (64.4\% v.s. 48.9\%). On the full test set, our system achieves 77\% accuracy, while off-the-shelf frontier LMs like o3 perform no better than random guessing, even with the same retrieval augmentation. We verify that our system does not exploit superficial features like idea complexity through extensive human-written and LM-designed robustness tests. Finally, we evaluate our system on unpublished novel ideas, including ideas generated by an AI ideation agent. Our system achieves 63.6\% accuracy, demonstrating its potential as a reward model for improving idea generation models. Altogether, our results outline a promising new direction for LMs to accelerate empirical AI research.
Jiaxin Wen, Chenglei Si, Yueh-Han Chen, He He 0001, Shi Feng 0005
NeurIPS5
2024 A SMART Mnemonic Sounds like "Glue Tonic": Mixing LLMs with Student Feedback to Make Mnemonic Learning Stick
abstract
Nishant Balepur, Matthew Shu, Alexander Hoyle, Alison Robey, Shi Feng, Seraphina Goldfarb-Tarrant, Jordan Lee Boyd-Graber. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Nishant Balepur, Matthew Shu, Alexander Miserlis Hoyle, Alison Robey, Shi Feng 0005, Seraphina Goldfarb-Tarrant, Jordan L. Boyd-Graber
EMNLP5
2024 KARL: Knowledge-Aware Retrieval and Representations aid Retention and Learning in Students
abstract
Flashcard schedulers rely on 1) student models to predict the flashcards a student knows; and 2) teaching policies to pick which cards to show next via these predictions.Prior student models, however, just use study data like the student's past responses, ignoring the text on cards.We propose content-aware scheduling, the first schedulers exploiting flashcard content.To give the first evidence that such schedulers enhance student learning, we build KAR 3 L, a simple but effective content-aware student model employing deep knowledge tracing (DKT), retrieval, and BERT to predict student recall.We train KAR 3 L by collecting a new dataset of 123,143 study logs on diverse trivia questions.KAR 3 L bests existing student models in AUC and calibration error.To ensure our improved predictions lead to better student learning, we create a novel delta-based teaching policy to deploy KAR 3 L online.Based on 32 study paths from 27 users, KAR 3 L improves learning efficiency over SOTA, showing KAR 3 L's strength and encouraging researchers to look beyond historical study data to fully capture student abilities.1 * Equal contribution.
Matthew Shu, Nishant Balepur, Shi Feng 0005, Jordan L. Boyd-Graber
EMNLP3
2024 Large Language Models Help Humans Verify Truthfulness - Except When They Are Convincingly Wrong
abstract
Chenglei Si, Navita Goyal, Tongshuang Wu, Chen Zhao, Shi Feng, Hal Daumé Iii, Jordan Boyd-Graber. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Chenglei Si, Navita Goyal, Sherry Tongshuang Wu, Chen Zhao 0013, Shi Feng 0005, Hal Daumé III, Jordan L. Boyd-Graber
NAACL-HLT5
2024 LLM Evaluators Recognize and Favor Their Own Generations
abstract
Self-evaluation using large language models (LLMs) has proven valuable not only in benchmarking but also methods like reward modeling, constitutional AI, and self-refinement. But new biases are introduced due to the same LLM acting as both the evaluator and the evaluatee. One such bias is self-preference, where an LLM evaluator scores its own outputs higher than others’ while human annotators consider them of equal quality. But do LLMs actually recognize their own outputs when they give those texts higher scores, or is it just a coincidence? In this paper, we investigate if self-recognition capability contributes to self-preference. We discover that, out of the box, LLMs such as GPT-4 and Llama 2 have non-trivial accuracy at distinguishing themselves from other LLMs and humans. By finetuning LLMs, we discover a linear correlation between self-recognition capability and the strength of self-preference bias; using controlled experiments, we show that the causal explanation resists straightforward confounders. We discuss how self-recognition can interfere with unbiased evaluations and AI safety more generally.
Arjun Panickssery, Samuel R. Bowman, Shi Feng 0005
NeurIPS3
2023 Measuring Inductive Biases of In-Context Learning with Underspecified Demonstrations
abstract
In-context learning (ICL) is an important paradigm for adapting large language models (LLMs) to new tasks, but the generalization behavior of ICL remains poorly understood.We investigate the inductive biases of ICL from the perspective of feature bias: which feature ICL is more likely to use given a set of underspecified demonstrations in which two features are equally predictive of the labels.First, we characterize the feature biases of GPT-3 models by constructing underspecified demonstrations from a range of NLP datasets and feature combinations.We find that LLMs exhibit clear feature biases-for example, demonstrating a strong bias to predict labels according to sentiment rather than shallow lexical features, like punctuation.Second, we evaluate the effect of different interventions that are designed to impose an inductive bias in favor of a particular feature, such as adding a natural language instruction or using semantically relevant label words.We find that, while many interventions can influence the learner to prefer a particular feature, it can be difficult to overcome strong prior biases.Overall, our results provide a broader picture of the types of features that ICL may be more likely to exploit and how to impose inductive biases that are better aligned with the intended task. 1
Chenglei Si, Dan Friedman, Nitish Joshi, Shi Feng 0005, Danqi Chen 0001, He He 0001
ACL (1)4
2023 Learning Human-Compatible Representations for Case-Based Decision Support
Yizhou Tian, Chacha Chen, Shi Feng 0005, Yuxin Chen 0001, Chenhao Tan
ICLR4
2022 Learning to Explain Selectively: A Case Study on Question Answering
abstract
Explanations promise to bridge the gap between humans and AI, yet it remains difficult to achieve consistent improvement in AIaugmented human decision making.The usefulness of AI explanations depends on many factors, and always showing the same type of explanation in all cases is suboptimal-so is relying on heuristics to adapt explanations for each scenario.We propose learning to explain "selectively": for each decision that the user makes, we use a model to choose the best explanation from a set of candidates and update this model with feedback to optimize human performance.We experiment on a question answering task, Quizbowl, and show that selective explanations improve human performance for both experts and crowdworkers.
Shi Feng 0005, Jordan L. Boyd-Graber
EMNLP1
2022 Active Example Selection for In-Context Learning
abstract
With a handful of demonstration examples, large-scale language models show strong capability to perform various tasks by in-context learning from these examples, without any finetuning.We demonstrate that in-context learning performance can be highly unstable across samples of examples, indicating the idiosyncrasies of how language models acquire information.We formulate example selection for in-context learning as a sequential decision problem, and propose a reinforcement learning algorithm for identifying generalizable policies to select demonstration examples.For GPT-2, our learned policies demonstrate strong abilities of generalizing to unseen tasks in training, with a 5.8% improvement on average.Examples selected from our learned policies can even achieve a small improvement on GPT-3 Ada.However, the improvement diminishes on larger GPT-3 models, suggesting emerging capabilities of large language models.
Yiming Zhang 0022, Shi Feng 0005, Chenhao Tan
EMNLP2
2021 Calibrate Before Use: Improving Few-shot Performance of Language Models
abstract
GPT-3 can perform numerous tasks when provided a natural language prompt that contains a few training examples. We show that this type of few-shot learning can be unstable: the choice of prompt format, training examples, and even the order of the examples can cause accuracy to vary from near chance to near state-of-the-art. We demonstrate that this instability arises from the bias of language models towards predicting certain answers, e.g., those that are placed near the end of the prompt or are common in the pre-training data. To mitigate this, we first estimate the model’s bias towards each answer by asking for its prediction when given a training prompt and a content-free test input such as "N/A". We then fit calibration parameters that cause the prediction for this input to be uniform across answers. On a diverse set of tasks, this contextual calibration procedure substantially improves GPT-3 and GPT-2’s accuracy (up to 30.0% absolute) across different choices of the prompt, while also making learning considerably more stable.
Eric Wallace, Shi Feng 0005, Daniel Klein 0001, Sameer Singh 0001
ICML3
2021 Concealed Data Poisoning Attacks on NLP Models
abstract
Adversarial attacks alter NLP model predictions by perturbing test-time inputs.However, it is much less understood whether, and how, predictions can be manipulated with small, concealed changes to the training data.In this work, we develop a new data poisoning attack that allows an adversary to control model predictions whenever a desired trigger phrase is present in the input.For instance, we insert 50 poison examples into a sentiment model's training set that causes the model to frequently predict Positive whenever the input contains "James Bond".Crucially, we craft these poison examples using a gradient-based procedure so that they do not mention the trigger phrase.We also apply our poison attack to language modeling ("Apple iPhone" triggers negative generations) and machine translation ("iced coffee" mistranslated as "hot coffee").We conclude by proposing three defenses that can mitigate our attack at some cost in prediction accuracy or extra human annotation.
Eric Wallace, Tony Z. Zhao, Shi Feng 0005, Sameer Singh 0001
NAACL-HLT3
2019 Misleading Failures of Partial-input Baselines
abstract
Recent work establishes dataset difficulty and removes annotation artifacts via partial-input baselines (e.g., hypothesis-only model for SNLI or question-only model for VQA). A successful partial-input baseline indicates that the dataset is cheatable. But the converse is not necessarily true: failures of partial-input baselines do not mean the dataset is free of artifacts. We first design artificial datasets to illustrate how the trivial patterns that are only visible in the full input can evade any partial-input baseline. Next, we identify such artifacts in the SNLI dataset—a hypothesis-only model augmented with trivial patterns in the premise can solve 15% of previously-thought “hard” examples. Our work provides a caveat for the use and creation of partial-input baselines for datasets.
Shi Feng 0005, Eric Wallace, Jordan L. Boyd-Graber
ACL (1)1
2019 Universal Adversarial Triggers for Attacking and Analyzing NLP
abstract
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, Sameer Singh. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Eric Wallace, Shi Feng 0005, Nikhil Kandpal, Matt Gardner 0001, Sameer Singh 0001
EMNLP/IJCNLP (1)2
2019 Understanding Impacts of High-Order Loss Approximations and Features in Deep Learning Interpretation
abstract
Current saliency map interpretations for neural networks generally rely on two key assumptions. First, they use first-order approximations of the loss function, neglecting higher-order terms such as the loss curvature. Second, they evaluate each feature’s importance in isolation, ignoring feature interdependencies. This work studies the effect of relaxing these two assumptions. First, we characterize a closed-form formula for the input Hessian matrix of a deep ReLU network. Using this formula, we show that, for classification problems with many classes, if a prediction has high probability then including the Hessian term has a small impact on the interpretation. We prove this result by demonstrating that these conditions cause the Hessian matrix to be approximately rank one and its leading eigenvector to be almost parallel to the gradient of the loss. We empirically validate this theory by interpreting ImageNet classifiers. Second, we incorporate feature interdependencies by calculating the importance of group-features using a sparsity regularization term. We use an L0 - L1 relaxation technique along with proximal gradient descent to efficiently compute group-feature importance values. Our empirical results show that our method significantly improves deep learning interpretations.
Sahil Singla 0002, Eric Wallace, Shi Feng 0005, Soheil Feizi
ICML3
2019 What can AI do for me?: evaluating machine learning interpretations in cooperative play
abstract
Machine learning is an important tool for decision making, but its ethical and responsible application requires rigorous vetting of its interpretability and utility: an understudied problem, particularly for natural language processing models. We propose an evaluation of interpretation on a real task with real human users, where the effectiveness of interpretation is measured by how much it improves human performance. We design a grounded, realistic human-computer cooperative setting using a question answering task, Quizbowl. We recruit both trivia experts and novices to play this game with computer as their teammate, who communicates its prediction via three different interpretations. We also provide design guidance for natural language processing human-in-the-loop settings.
Shi Feng 0005, Jordan L. Boyd-Graber
IUI1
2019 Trick Me If You Can: Human-in-the-loop Generation of Adversarial Question Answering Examples
abstract
Adversarial evaluation stress-tests a model’s understanding of natural language. Because past approaches expose superficial patterns, the resulting adversarial examples are limited in complexity and diversity. We propose human- in-the-loop adversarial generation, where human authors are guided to break models. We aid the authors with interpretations of model predictions through an interactive user interface. We apply this generation framework to a question answering task called Quizbowl, where trivia enthusiasts craft adversarial questions. The resulting questions are validated via live human–computer matches: Although the questions appear ordinary to humans, they systematically stump neural and information retrieval models. The adversarial questions cover diverse phenomena from multi-hop reasoning to entity type distractors, exposing open challenges in robust question answering.
Eric Wallace, Pedro Rodríguez 0001, Shi Feng 0005, Ikuya Yamada, Jordan L. Boyd-Graber
Trans. Assoc. Comput. Linguistics3
2018 Pathologies of Neural Models Make Interpretation Difficult
abstract
One way to interpret neural model predictions is to highlight the most important input features-for example, a heatmap visualization over the words in an input sentence.In existing interpretation methods for NLP, a word's importance is determined by either input perturbation-measuring the decrease in model confidence when that word is removed-or by the gradient with respect to that word.To understand the limitations of these methods, we use input reduction, which iteratively removes the least important word from the input.This exposes pathological behaviors of neural models: the remaining words appear nonsensical to humans and are not the ones determined as important by interpretation methods.As we confirm with human experiments, the reduced examples lack information to support the prediction of any label, but models still make the same predictions with high confidence.To explain these counterintuitive results, we draw connections to adversarial examples and confidence calibration: pathological behaviors reveal difficulties in interpreting neural models trained with maximum likelihood.To mitigate their deficiencies, we fine-tune the models by encouraging high entropy outputs on reduced examples.Fine-tuned models become more interpretable under input reduction without accuracy loss on regular examples.
Shi Feng 0005, Eric Wallace, Alvin Grissom II, Mohit Iyyer, Pedro Rodríguez 0001, Jordan L. Boyd-Graber
EMNLP1