Emmy Liu

dblp:249/6997 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
11since 2021 · last 2025
0000-0003-1287-5712ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 6 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Not-Just-Scaling Laws: Towards a Better Understanding of the Downstream Impact of Language Model Design Decisions
abstract
Emmy Liu, Amanda Bertsch, Lintang Sutawika, Lindia Tjuatja, Patrick Fernandes, Lara Marinov, Michael Chen, Shreya Singhal, Carolin Lawrence, Aditi Raghunathan, Kiril Gashteovski, Graham Neubig. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Emmy Liu, Amanda Bertsch, Lintang Sutawika, Lindia Tjuatja, Patrick Fernandes, Lara Marinov, Shreya Singhal, Carolin Lawrence, Aditi Raghunathan, Kiril Gashteovski, Graham Neubig
EMNLP1
2024 Program-Aided Reasoners (Better) Know What They Know
abstract
Anubha Kabra, Sanketh Rangreji, Yash Mathur, Aman Madaan, Emmy Liu, Graham Neubig. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Anubha Kabra, Sanketh Rangreji, Yash Mathur, Aman Madaan, Emmy Liu, Graham Neubig
NAACL-HLT5
2024 Divergences between Language Models and Human Brains
abstract
Do machines and humans process language in similar ways? Recent research has hinted at the affirmative, showing that human neural activity can be effectively predicted using the internal representations of language models (LMs). Although such results are thought to reflect shared computational principles between LMs and human brains, there are also clear differences in how LMs and humans represent and use language. In this work, we systematically explore the divergences between human and machine language processing by examining the differences between LM representations and human brain responses to language as measured by Magnetoencephalography (MEG) across two datasets in which subjects read and listened to narrative stories. Using an LLM-based data-driven approach, we identify two domains that LMs do not capture well: social/emotional intelligence and physical commonsense. We validate these findings with human behavioral experiments and hypothesize that the gap is due to insufficient representations of social/emotional and physical knowledge in LMs. Our results show that fine-tuning LMs on these domains can improve their alignment with human brain responses.
Yuchen Zhou 0004, Emmy Liu, Graham Neubig, Michael J. Tarr, Leila Wehbe
NeurIPS2
2023 When Does Translation Require Context? A Data-driven, Multilingual Exploration
abstract
Although proper handling of discourse significantly contributes to the quality of machine translation (MT), these improvements are not adequately measured in common translation quality metrics.Recent works in context-aware MT attempt to target a small set of discourse phenomena during evaluation, however not in a fully systematic way.In this paper, we develop the Multilingual Discourse-Aware (MUDA) benchmark, a series of taggers that identify and evaluate model performance on discourse phenomena in any given dataset.The choice of phenomena is inspired by a novel methodology to systematically identify translations requiring context.We confirm the difficulty of previously studied phenomena while uncovering others that were previously unaddressed.We find that common context-aware MT models make only marginal improvements over context-agnostic models, which suggests these models do not handle these ambiguities effectively.We release code and data for 14 language pairs to encourage the MT community to focus on accurately capturing discourse phenomena.1
Patrick Fernandes, Kayo Yin, Emmy Liu, André F. T. Martins, Graham Neubig
ACL (1)3
2023 Exam Eustress: Designing Brief Online Interventions for Helping Students Identify Positive Aspects of Stress
abstract
Stress reappraisal interventions try to shift students’ negative perceptions towards eustress, stress that can be beneficial, and help them perform better. However, it is less clear how to present them to users as online interventions that are brief, voluntary, and scale well in real-world contexts. We explore the design of online exam eustress interventions by generating six design factors (D1-6) that reinforce a core reappraisal message (D0), and evaluate them through: (i) user interviews (N = 20) revealing six findings (F1-6) on the importance of elaboration, layout, modality, and source of intervention content; (ii) a field experiment (N = 1283) showing a significant positive effect on exam scores (p = 0.003). Subgroup analyses indicate a significant effect for first-year but not for upper-year students, and no detectable gender differences. Our work offers insight into how students interact with online mindset interventions and design considerations for incorporating them into large courses.
Mohi Reza, Angela M. Zavaleta Bernuy, Emmy Liu, Zhongyuan Liang, Calista K. Barber, Joseph Jay Williams
CHI3
2023 Crossing the Threshold: Idiomatic Machine Translation through Retrieval Augmentation and Loss Weighting
abstract
Idioms are common in everyday language, but often pose a challenge to translators because their meanings do not follow from the meanings of their parts.Despite significant advances, machine translation systems still struggle to translate idiomatic expressions.We provide a simple characterization of idiomatic translation and related issues.This allows us to conduct a synthetic experiment revealing a tipping point at which transformer-based machine translation models correctly default to idiomatic translations.To expand multilingual resources, we compile a dataset of ∼ 4k natural sentences containing idiomatic expressions in French, Finnish, and Japanese.To improve translation of natural idioms, we introduce two straightforward yet effective techniques: the strategic upweighting of training loss on potentially idiomatic sentences, and using retrievalaugmented models.This not only improves the accuracy of a strong pretrained MT model on idiomatic sentences by up to 13% in absolute accuracy, but also holds potential benefits for non-idiomatic sentences.1
Emmy Liu, Aditi Chaudhary, Graham Neubig
EMNLP1
2023 Computational Language Acquisition with Theory of Mind
Andy Liu, Hao Zhu 0011, Emmy Liu, Yonatan Bisk, Graham Neubig
ICLR3
2023 Bridging the Gap: A Survey on Integrating (Human) Feedback for Natural Language Generation
abstract
Abstract Natural language generation has witnessed significant advancements due to the training of large language models on vast internet-scale datasets. Despite these advancements, there exists a critical challenge: These models can inadvertently generate content that is toxic, inaccurate, and unhelpful, and existing automatic evaluation metrics often fall short of identifying these shortcomings. As models become more capable, human feedback is an invaluable signal for evaluating and improving models. This survey aims to provide an overview of recent research that has leveraged human feedback to improve natural language generation. First, we introduce a taxonomy distilled from existing research to categorize and organize the varied forms of feedback. Next, we discuss how feedback can be described by its format and objective, and cover the two approaches proposed to use feedback (either for training or decoding): directly using feedback or training feedback models. We also discuss existing datasets for human-feedback data collection, and concerns surrounding feedback collection. Finally, we provide an overview of the nascent field of AI feedback, which uses large language models to make judgments based on a set of principles and minimize the need for human intervention. We also release a website of this survey at feedback-gap-survey.info.
Patrick Fernandes, Aman Madaan, Emmy Liu, António Farinhas, Pedro Henrique Martins, Amanda Bertsch, José Guilherme Camargo de Souza, Shuyan Zhou, Sherry Tongshuang Wu, Graham Neubig, André F. T. Martins
Trans. Assoc. Comput. Linguistics3
2022 The emergence of moral foundations in child language development
Aida Ramezani, Emmy Liu, Renato Ferreira Pinto Junior, Spike W. S. Lee, Yang Xu 0023
CogSci2
2022 Are representations built from the ground up? An empirical examination of local composition in language models
abstract
Compositionality, the phenomenon where the meaning of a phrase can be derived from its constituent parts, is a hallmark of human language.At the same time, many phrases are non-compositional, carrying a meaning beyond that of each part in isolation.Representing both of these types of phrases is critical for language understanding, but it is an open question whether modern language models (LMs) learn to do so; in this work we examine this question.We first formulate a problem of predicting the LM-internal representations of longer phrases given those of their constituents.We find that the representation of a parent phrase can be predicted with some accuracy given an affine transformation of its children.While we would expect the predictive accuracy to correlate with human judgments of semantic compositionality, we find this is largely not the case, indicating that LMs may not accurately distinguish between compositional and non-compositional phrases.We perform a variety of analyses, shedding light on when different varieties of LMs do and do not generate compositional representations, and discuss implications for future modeling work. 1
Emmy Liu, Graham Neubig
EMNLP1
2022 Testing the Ability of Language Models to Interpret Figurative Language
abstract
Emmy Liu, Chenxuan Cui, Kenneth Zheng, Graham Neubig. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Emmy Liu, Chenxuan Cui, Kenneth Zheng, Graham Neubig
NAACL-HLT1
2020 Chaining and the process of scientific innovation
Emmy Liu, Yang Xu 0023
CogSci1
2019 Rapid information gain explains cross-linguistic tendencies in numeral ordering
Emmy Liu, Yang Xu 0023
CogSci1