Angelica Chen

dblp:241/5892 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
9since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 9 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Language models and text generation · 40% Reinforcement learning · 11% Efficient and distributed learning · 9%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 82% Computing education · 18%

Topics — the 23 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
alignment
1.422024
Preference Learning Algorithms Do Not Learn Preference Rankings · NeurIPS 2024
Pretraining Language Models with Human Preferences · ICML 2023
Machine learning › Optimization for machine learning
bilevel optimization
0.912025
Generalists vs. Specialists: Evaluating LLMs on Highly-Constrained Biophysical Sequence Optimization Tasks · ICML 2025
Natural language and speech › Language models and text generation › large language model
LLM-based optimization
0.912025
Generalists vs. Specialists: Evaluating LLMs on Highly-Constrained Biophysical Sequence Optimization Tasks · ICML 2025
Machine learning › Trustworthy machine learning
interpretability
0.812024
Sudden Drops in the Loss: Syntax Acquisition, Phase Transitions, and Simplicity Bias in MLMs · ICLR 2024
Natural language and speech › Language models and text generation
masked language modeling
0.812024
Sudden Drops in the Loss: Syntax Acquisition, Phase Transitions, and Simplicity Bias in MLMs · ICLR 2024
Machine learning › Reinforcement learning
preference learning
0.812024
Preference Learning Algorithms Do Not Learn Preference Rankings · NeurIPS 2024
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.812024
Preference Learning Algorithms Do Not Learn Preference Rankings · NeurIPS 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning
syntax acquisition
0.812024
Sudden Drops in the Loss: Syntax Acquisition, Phase Transitions, and Simplicity Bias in MLMs · ICLR 2024
Machine learning › Deep learning architectures and training
training dynamics
0.812024
Sudden Drops in the Loss: Syntax Acquisition, Phase Transitions, and Simplicity Bias in MLMs · ICLR 2024
Natural language and speech › Language models and text generation
code generation
0.712023
EvoPrompting: Language Models for Code-Level Neural Architecture Search · NeurIPS 2023
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search
evolutionary neural architecture search
0.712023
EvoPrompting: Language Models for Code-Level Neural Architecture Search · NeurIPS 2023
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search
0.712023
EvoPrompting: Language Models for Code-Level Neural Architecture Search · NeurIPS 2023
Machine learning › Representation and self-supervised learning
pre-training
0.712023
Pretraining Language Models with Human Preferences · ICML 2023
Natural language and speech › Language models and text generation › text summarization
long document summarization
0.612022
SQuALITY: Building a Long-Document Summarization Dataset the Hard Way · EMNLP 2022
Natural language and speech › Language models and text generation › text summarization › controllable summarization
query-focused summarization
0.612022
SQuALITY: Building a Long-Document Summarization Dataset the Hard Way · EMNLP 2022
Natural language and speech › Language models and text generation
text summarization
0.612022
SQuALITY: Building a Long-Document Summarization Dataset the Hard Way · EMNLP 2022
Natural language and speech › Question answering and dialogue systems › knowledge base question answering
logical form generation
0.412019
Generating Logical Forms from Graph Representations of Text and Entities · ACL (1) 2019
Natural language and speech › Information extraction and text analysis
semantic parsing
0.412019
Generating Logical Forms from Graph Representations of Text and Entities · ACL (1) 2019
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization
0.212024
Preference Learning Algorithms Do Not Learn Preference Rankings · NeurIPS 2024
Machine learning › Learning theory
phase transition
0.212024
Sudden Drops in the Loss: Syntax Acquisition, Phase Transitions, and Simplicity Bias in MLMs · ICLR 2024
Information retrieval
evaluation
0.212022
SQuALITY: Building a Long-Document Summarization Dataset the Hard Way · EMNLP 2022
Information retrieval › text summarization
summarization evaluation
0.212022
SQuALITY: Building a Long-Document Summarization Dataset the Hard Way · EMNLP 2022
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph
0.112019
Generating Logical Forms from Graph Representations of Text and Entities · ACL (1) 2019

Methods — techniques the papers use, named apart from their topics

preference learning · 1.7black-box optimization · 1.7bi-level optimization · 1.7survey · 1.3ranking accuracy analysis · 0.8idealized ranking accuracy · 0.8causal manipulation · 0.8attention analysis · 0.8reinforcement learning from human feedback · 0.7conditional training · 0.7human annotation · 0.6automatic evaluation metrics · 0.6
YearPublicationVenuePosition
2025 Generalists vs. Specialists: Evaluating LLMs on Highly-Constrained Biophysical Sequence Optimization Tasks
abstract
Although large language models (LLMs) have shown promise in biomolecule optimization problems, they incur heavy computational costs and struggle to satisfy precise constraints. On the other hand, specialized solvers like LaMBO-2 offer efficiency and fine-grained control but require more domain expertise. Comparing these approaches is challenging due to expensive laboratory validation and inadequate synthetic benchmarks. We address this by introducing Ehrlich functions, a synthetic test suite that captures the geometric structure of biophysical sequence optimization problems. With prompting alone, off-the-shelf LLMs struggle to optimize Ehrlich functions. In response, we propose LLOME (Language Model Optimization with Margin Expectation), a bilevel optimization routine for online black-box optimization. When combined with a novel preference learning loss, we find LLOME can not only learn to solve some Ehrlich functions, but can even perform as well as or better than LaMBO-2 on moderately difficult Ehrlich variants. However, LLMs also exhibit some likelihood-reward miscalibration and struggle without explicit rewards. Our results indicate LLMs can occasionally provide significant benefits, but specialized solvers are still competitive and incur less overhead.
Angelica Chen, Samuel Stanton, Frances Ding, Robert G. Alberstein, Andrew M. Watkins, Richard Bonneau, Vladimir Gligorijevic, Kyunghyun Cho, Nathan C. Frey
ICML1
2024 Sudden Drops in the Loss: Syntax Acquisition, Phase Transitions, and Simplicity Bias in MLMs
abstract
Most interpretability research in NLP focuses on understanding the behavior and features of a fully trained model. However, certain insights into model behavior may only be accessible by observing the trajectory of the training process. We present a case study of syntax acquisition in masked language models (MLMs) that demonstrates how analyzing the evolution of interpretable artifacts throughout training deepens our understanding of emergent behavior. In particular, we study Syntactic Attention Structure (SAS), a naturally emerging property of MLMs wherein specific Transformer heads tend to focus on specific syntactic relations. We identify a brief window in pretraining when models abruptly acquire SAS, concurrent with a steep drop in loss. This breakthrough precipitates the subsequent acquisition of linguistic capabilities. We then examine the causal role of SAS by manipulating SAS during training, and demonstrate that SAS is necessary for the development of grammatical capabilities. We further find that SAS competes with other beneficial traits during training, and that briefly suppressing SAS improves model quality. These findings offer an interpretation of a real-world example of both simplicity bias and breakthrough training dynamics.
Angelica Chen, Ravid Shwartz-Ziv, Kyunghyun Cho, Matthew L. Leavitt, Naomi Saphra
ICLR1
2024 Preference Learning Algorithms Do Not Learn Preference Rankings
abstract
Preference learning algorithms (e.g., RLHF and DPO) are frequently used to steer LLMs to produce generations that are more preferred by humans, but our understanding of their inner workings is still limited. In this work, we study the conventional wisdom that preference learning trains models to assign higher likelihoods to more preferred outputs than less preferred outputs, measured via *ranking accuracy*. Surprisingly, we find that most state-of-the-art preference-tuned models achieve a ranking accuracy of less than 60% on common preference datasets. We furthermore derive the *idealized ranking accuracy* that a preference-tuned LLM would achieve if it optimized the DPO or RLHF objective perfectly. We demonstrate that existing models exhibit a significant *alignment gap* -- *i.e.*, a gap between the observed and idealized ranking accuracies. We attribute this discrepancy to the DPO objective, which is empirically and theoretically ill-suited to correct even mild ranking errors in the reference model, and derive a simple and efficient formula for quantifying the difficulty of learning a given preference datapoint. Finally, we demonstrate that ranking accuracy strongly correlates with the empirically popular win rate metric when the model is close to the reference model used in the objective, shedding further light on the differences between on-policy (e.g., RLHF) and off-policy (e.g., DPO) preference learning algorithms.
Angelica Chen, Sadhika Malladi, Lily H. Zhang, Xinyi Chen 0001, Qiuyi Zhang 0001, Rajesh Ranganath, Kyunghyun Cho
NeurIPS1
2023 What Do NLP Researchers Believe? Results of the NLP Community Metasurvey
abstract
Julian Michael, Ari Holtzman, Alicia Parrish, Aaron Mueller, Alex Wang, Angelica Chen, Divyam Madaan, Nikita Nangia, Richard Yuanzhe Pang, Jason Phang, Samuel R. Bowman. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Julian Michael, Ari Holtzman, Alicia Parrish, Aaron Mueller, Angelica Chen, Divyam Madaan, Nikita Nangia, Richard Yuanzhe Pang, Jason Phang, Samuel R. Bowman
ACL (1)6
2023 Pretraining Language Models with Human Preferences
abstract
Language models (LMs) are pretrained to imitate text from large and diverse datasets that contain content that would violate human preferences if generated by an LM: falsehoods, offensive comments, personally identifiable information, low-quality or buggy code, among others. Here, we explore alternative objectives for pretraining LMs in a way that also guides them to generate text aligned with human preferences. We benchmark five objectives for pretraining with human feedback across three tasks and study how they affect the alignment and capabilities of pretrained LMs. We find a Pareto-optimal and simple approach among those we explored: conditional training, or learning distribution over tokens conditional on their human preference scores. Conditional training reduces the rate of undesirable content by up to an order of magnitude, both when generating without a prompt and with an adversarially-chosen prompt. Moreover, conditional training maintains the downstream task performance of standard LM pretraining, both before and after task-specific finetuning. Pretraining with human feedback results in much better preference satisfaction than standard LM pretraining followed by finetuning with feedback, i.e., learning and then unlearning undesirable behavior. Our results suggest that we should move beyond imitation learning when pretraining LMs and incorporate human preferences from the start of training.
Tomasz Korbak, Kejian Shi, Angelica Chen, Rasika Bhalerao, Christopher L. Buckley, Jason Phang, Samuel R. Bowman, Ethan Perez
ICML3
2023 EvoPrompting: Language Models for Code-Level Neural Architecture Search
abstract
Given the recent impressive accomplishments of language models (LMs) for code generation, we explore the use of LMs as general adaptive mutation and crossover operators for an evolutionary neural architecture search (NAS) algorithm. While NAS still proves too difficult a task for LMs to succeed at solely through prompting, we find that the combination of evolutionary prompt engineering with soft prompt-tuning, a method we term EvoPrompting, consistently finds diverse and high performing models. We first demonstrate that EvoPrompting is effective on the computationally efficient MNIST-1D dataset, where EvoPrompting produces convolutional architecture variants that outperform both those designed by human experts and naive few-shot prompting in terms of accuracy and model size. We then apply our method to searching for graph neural networks on the CLRS Algorithmic Reasoning Benchmark, where EvoPrompting is able to design *novel* architectures that outperform current state-of-the-art models on 21 out of 30 algorithmic reasoning tasks while maintaining similar model size. EvoPrompting is successful at designing accurate and efficient neural network architectures across a variety of machine learning tasks, while also being general enough for easy adaptation to other tasks beyond neural network design.
Angelica Chen, David Dohan, David R. So
NeurIPS1
2022 SQuALITY: Building a Long-Document Summarization Dataset the Hard Way
abstract
Summarization datasets are often assembled either by scraping naturally occurring publicdomain summaries-which are nearly always in difcult-to-work-with technical domainsor by using approximate heuristics to extract them from everyday text-which frequently yields unfaithful summaries.In this work, we turn to a slower but more straightforward approach to developing summarization benchmark data: We hire highly-qualied contractors to read stories and write original summaries from scratch.To amortize reading time, we collect ve summaries per document, with the rst giving an overview and the subsequent four addressing specic questions.We use this protocol to collect SQuAL-ITY, a dataset of question-focused summaries built on the same public-domain short stories as the multiple-choice dataset QuALITY (Pang et al., 2021b).Experiments with stateof-the-art summarization systems show that our dataset is challenging and that existing automatic evaluation metrics are weak indicators of quality.
Richard Yuanzhe Pang, Angelica Chen, Jason Phang, Samuel R. Bowman
EMNLP3
2022 Teaching BERT to Wait: Balancing Accuracy and Latency for Streaming Disfluency Detection
abstract
Angelica Chen, Vicky Zayats, Daniel Walker, Dirk Padfield. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Angelica Chen, Victoria Zayats, Daniel D. Walker, Dirk Padfield
NAACL-HLT1
2022 QuALITY: Question Answering with Long Input Texts, Yes!
abstract
Richard Yuanzhe Pang, Alicia Parrish, Nitish Joshi, Nikita Nangia, Jason Phang, Angelica Chen, Vishakh Padmakumar, Johnny Ma, Jana Thompson, He He, Samuel Bowman. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Richard Yuanzhe Pang, Alicia Parrish, Nitish Joshi, Nikita Nangia, Jason Phang, Angelica Chen, Vishakh Padmakumar, Johnny Ma, Jana Thompson, He He 0001, Samuel R. Bowman
NAACL-HLT6
2019 Generating Logical Forms from Graph Representations of Text and Entities
abstract
Structured information about entities is critical for many semantic parsing tasks.We present an approach that uses a Graph Neural Network (GNN) architecture to incorporate information about relevant entities and their relations during parsing.Combined with a decoder copy mechanism, this approach provides a conceptually simple mechanism to generate logical forms with entities.We demonstrate that this approach is competitive with the stateof-the-art across several tasks without pretraining, and outperforms existing approaches when combined with BERT pre-training.
Peter Shaw 0004, Philip Massey, Angelica Chen, Francesco Piccinno, Yasemin Altun
ACL (1)3