Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Melanie Subbiah

dblp:266/3277 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 7 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Language models and text generation · 59% Trustworthy machine learning · 32% Transfer learning and domain adaptation · 4%

Topics — the 11 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › evaluation of language models
faithfulness evaluation
1.622025
Is the Top Still Spinning? Evaluating Subjectivity in Narrative Understanding · EMNLP 2025
STORYSUMM: Evaluating Faithfulness in Story Summarization · EMNLP 2024
Natural language and speech › Language models and text generation › text summarization › domain-specific summarization
narrative summarization
1.622025
Is the Top Still Spinning? Evaluating Subjectivity in Narrative Understanding · EMNLP 2025
STORYSUMM: Evaluating Faithfulness in Story Summarization · EMNLP 2024
Natural language and speech › Language models and text generation
text summarization
1.622025
Is the Top Still Spinning? Evaluating Subjectivity in Narrative Understanding · EMNLP 2025
STORYSUMM: Evaluating Faithfulness in Story Summarization · EMNLP 2024
Machine learning › Trustworthy machine learning
interpretability
0.922024
Unsupervised Selective Rationalization with Noise Injection · ACL (1) 2023
STORYSUMM: Evaluating Faithfulness in Story Summarization · EMNLP 2024
Machine learning › Trustworthy machine learning
fairness
0.912025
Guiding LLM Decision-Making with Fairness Reward Models · NeurIPS 2025
Natural language and speech › Language models and text generation › LLM agents
LLM decision-making
0.912025
Guiding LLM Decision-Making with Fairness Reward Models · NeurIPS 2025
Machine learning › Trustworthy machine learning › interpretability › rationalization
selective rationalization
0.712023
Unsupervised Selective Rationalization with Noise Injection · ACL (1) 2023
Natural language and speech › Language models and text generation › neural language model
autoregressive language model
0.412020
Language Models are Few-Shot Learners · NeurIPS 2020
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.412020
Language Models are Few-Shot Learners · NeurIPS 2020
Machine learning › Trustworthy machine learning › interpretability › model debugging
failure explanation
0.212024
STORYSUMM: Evaluating Faithfulness in Story Summarization · EMNLP 2024
Natural language and speech › Information extraction and text analysis
text classification
0.212023
Unsupervised Selective Rationalization with Noise Injection · ACL (1) 2023

Methods — techniques the papers use, named apart from their topics

large language model · 1.6weak supervision · 0.9summary editing · 0.9reward modeling · 0.9chain-of-thought sampling · 0.9automatic evaluation metric · 0.9automatic evaluation metrics · 0.8noise injection · 0.7joint training · 0.7benchmark evaluation · 0.6
YearPublicationVenuePosition
2025 Is the Top Still Spinning? Evaluating Subjectivity in Narrative Understanding
abstract
Determining faithfulness of a claim to a source document is an important problem across many domains.This task is generally treated as a binary judgment of whether the claim is supported or unsupported in relation to the source.In many cases, though, whether a claim is supported can be ambiguous.For instance, it may depend on making inferences from given evidence, and different people can reasonably interpret the claim as either supported or unsupported based on their agreement with those inferences.Forcing binary labels upon such claims lowers the reliability of evaluation.In this work, we reframe the task to manage the subjectivity involved with factuality judgments of ambiguous claims.We introduce LLMgenerated edits of summaries as a method of providing a nuanced evaluation of claims: how much does a summary need to be edited to be unambiguous?Whether a claim gets rewritten and how much it changes can be used as an automatic evaluation metric, the Ambiguity Rewrite Metric (ARM), with a much richer feedback signal than a binary judgment of faithfulness.We focus on the area of narrative summarization as it is particularly rife with ambiguity and subjective interpretation.We show that ARM produces a 21% absolute improvement in annotator agreement on claim faithfulness, indicating that subjectivity is reduced.
Melanie Subbiah, Akankshya Mishra, Grace Kim, Liyan Tang, Greg Durrett, Kathy McKeown
EMNLP1
2025 Counterfactual Simulatability of LLM Explanations for Generation Tasks
abstract
LLMs can be unpredictable, as even slight alterations to the prompt can cause the output to change in unexpected ways. Thus, the ability of models to accurately explain their behavior is critical, especially in high-stakes settings. Counterfactual simulatability measures how well an explanation allows users to infer the model’s output on related counterfactuals and has been previously studied for yes/no question answering. We provide a general framework for extending this method to generation tasks, using news summarization and medical suggestion as example use cases. We find that while LLM explanations do enable users to better predict their outputs on counterfactuals in the summarization setting, there is significant room for improvement for medical suggestion. Furthermore, our results suggest that evaluating counterfactual simulatability may be more appropriate for skill-based tasks as opposed to knowledge-based tasks.
Marvin Limpijankit, Yanda Chen, Melanie Subbiah, Nicholas Deas, Kathy McKeown
INLG3
2025 Guiding LLM Decision-Making with Fairness Reward Models
abstract
Large language models are increasingly used to support high-stakes decisions, potentially influencing who is granted bail or receives a loan. Naive chain-of-thought sampling can improve average decision accuracy, but has also been shown to amplify unfair bias. To address this challenge and enable the trustworthy use of reasoning models in high-stakes decision-making, we propose a framework for training a generalizable Fairness Reward Model (FRM). Our model assigns a fairness score to LLM reasoning, enabling the system to down-weight biased trajectories and favor equitable ones when aggregating decisions across reasoning chains. We show that a single Fairness Reward Model, trained on weakly supervised, LLM-annotated examples of biased versus unbiased reasoning, transfers across tasks, domains, and model families without additional fine-tuning. When applied to real-world decision-making tasks including recidivism prediction and social media moderation, our approach consistently improves fairness while matching, or even surpassing, baseline accuracy.
Zara Hall, Melanie Subbiah, Thomas P. Zollo, Kathy McKeown, Richard S. Zemel
NeurIPS2
2024 STORYSUMM: Evaluating Faithfulness in Story Summarization
abstract
Human evaluation has been the gold standard for checking faithfulness in abstractive summarization.However, with a challenging source domain like narrative, multiple annotators can agree a summary is faithful, while missing details that are obvious errors only once pointed out.We therefore introduce a new dataset, STORYSUMM, comprising LLM summaries of short stories with localized faithfulness labels and error explanations.This benchmark is for evaluation methods, testing whether a given method can detect challenging inconsistencies.Using this dataset, we first show that any one human annotation protocol is likely to miss inconsistencies, and we advocate for pursuing a range of methods when establishing ground truth for a summarization dataset.We finally test recent automatic metrics and find that none of them achieve more than 70% balanced accuracy on this task, demonstrating that it is a challenging benchmark for future work in faithfulness evaluation.
Melanie Subbiah, Faisal Ladhak, Akankshya Mishra, Griffin Adams, Lydia B. Chilton, Kathy McKeown
EMNLP1
2024 Reading Subtext: Evaluating Large Language Models on Short Story Summarization with Writers
abstract
Abstract We evaluate recent Large Language Models (LLMs) on the challenging task of summarizing short stories, which can be lengthy, and include nuanced subtext or scrambled timelines. Importantly, we work directly with authors to ensure that the stories have not been shared online (and therefore are unseen by the models), and to obtain informed evaluations of summary quality using judgments from the authors themselves. Through quantitative and qualitative analysis grounded in narrative theory, we compare GPT-4, Claude-2.1, and LLama-2-70B. We find that all three models make faithfulness mistakes in over 50% of summaries and struggle with specificity and interpretation of difficult subtext. We additionally demonstrate that LLM ratings and other automatic metrics for summary quality do not correlate well with the quality ratings from the writers.
Melanie Subbiah, Sean Zhang, Lydia B. Chilton, Kathy McKeown
Trans. Assoc. Comput. Linguistics1
2023 Unsupervised Selective Rationalization with Noise Injection
abstract
A major issue with using deep learning models in sensitive applications is that they provide no explanation for their output.To address this problem, unsupervised selective rationalization produces rationales alongside predictions by chaining two jointly-trained components, a rationale generator and a predictor.Although this architecture guarantees that the prediction relies solely on the rationale, it does not ensure that the rationale contains a plausible explanation for the prediction.We introduce a novel training technique that effectively limits generation of implausible rationales by injecting noise between the generator and the predictor.Furthermore, we propose a new benchmark for evaluating unsupervised selective rationalization models using movie reviews from existing datasets.We achieve sizeable improvements in rationale plausibility and task accuracy over the state-of-the-art across a variety of tasks, including our new benchmark, while maintaining or improving model faithfulness.1
Adam Storek, Melanie Subbiah, Kathy McKeown
ACL (1)2
2022 SafeText: A Benchmark for Exploring Physical Safety in Language Models
abstract
Sharon Levy, Emily Allaway, Melanie Subbiah, Lydia Chilton, Desmond Patton, Kathleen McKeown, William Yang Wang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Sharon Levy, Emily Allaway, Melanie Subbiah, Lydia B. Chilton, Desmond Upton Patton, Kathy McKeown, William Yang Wang
EMNLP3
2020 Language Models are Few-Shot Learners
abstract
We demonstrate that scaling up language models greatly improves task-agnostic, few-shot performance, sometimes even becoming competitive with prior state-of-the-art fine-tuning approaches. Specifically, we train GPT-3, an autoregressive language model with 175 billion parameters, 10x more than any previous non-sparse language model, and test its performance in the few-shot setting. For all tasks, GPT-3 is applied without any gradient updates or fine-tuning, with tasks and few-shot demonstrations specified purely via text interaction with the model. GPT-3 achieves strong performance on many NLP datasets, including translation, question-answering, and cloze tasks. We also identify some datasets where GPT-3's few-shot learning still struggles, as well as some datasets where GPT-3 faces methodological issues related to training on large web corpora.
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Thomas Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu 0003, Clemens Winter, Christopher Hesse, Mark Chen 0003, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, Dario Amodei
NeurIPS4