Najoung Kim

dblp:194/1249 · DBLP profile ↗
← Back
25ranked-venue papers
7as first author
19since 2021 · last 2026
0009-0009-7352-8643ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 7 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 RExBench: Can coding agents autonomously implement AI research extensions?
abstract
Nicholas Edwards, Yukyung Lee, Yujun Audrey Mao, Yulu Qin, Sebastian Schuster, Najoung Kim. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Nicholas Edwards, Yukyung Lee, Yujun Audrey Mao, Yulu Qin, Sebastian Schuster 0001, Najoung Kim
ACL (1)6
2026 Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean Dialogues
abstract
Eunsu Kim, Junyeong Park, Juhyun Oh, Kiwoong Park, Seyoung Song, A. Seza Doğruöz, Alice Oh, Najoung Kim. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Eunsu Kim, Junyeong Park, Juhyun Oh, Kiwoong Park, Seyoung Song 0001, A. Seza Dogruöz, Alice Oh, Najoung Kim
ACL (1)8
2025 CheckEval: A reliable LLM-as-a-Judge framework for evaluating text generation using checklists
abstract
Existing LLM-as-a-Judge approaches for evaluating text generation suffer from rating inconsistencies, with low agreement and high rating variance across different evaluator models.We attribute this to subjective evaluation criteria combined with Likert scale scoring in existing protocols.To address this issue, we introduce CheckEval, a checklist-based evaluation framework that improves rating reliability via decomposed binary questions.Through experiments with 12 evaluator models across multiple datasets, we first demonstrate that CheckEval strongly correlates with human judgments.More importantly, CheckEval dramatically improves the average agreement across evaluator models by 0.45 and reduces the score variance.CheckEval scores furthermore have the benefit of being more interpretable because it decomposes evaluation criteria into traceable binary decisions, allowing analyses of specific attributes driving quality judgments.
Yukyung Lee, JoongHoon Kim, Jaehee Kim, Hyowon Cho, Jaewook Kang, Pilsung Kang 0001, Najoung Kim
EMNLP7
2025 Erasing More Than Intended? How Concept Erasure Degrades the Generation of Non-Target Concepts
Ibtihel Amara, Ahmed Imtiaz Humayun, Ivana Kajic, Zarana Parekh, Natalie Harris, Sarah Young, Chirag Nagpal, Najoung Kim, Junfeng He, Cristina Nader Vasconcelos, Deepak Ramachandran, Golnoosh Farnadi, Katherine A. Heller, Mohammad Havaei, Negar Rostamzadeh
ICCV8
2025 Transformers Struggle to Learn to Search
abstract
Search is an ability foundational in many important tasks, and recent studies have shown that large language models (LLMs) struggle to perform search robustly. It is unknown whether this inability is due to a lack of data, insufficient model parameters, or fundamental limitations of the transformer architecture. In this work, we use the foundational graph connectivity problem as a testbed to generate effectively limitless high-coverage data to train small transformers and test whether they can learn to perform search. We find that, when given the right training distribution, the transformer is able to learn to search. We analyze the algorithm that the transformer has learned through a novel mechanistic interpretability technique that enables us to extract the computation graph from the trained model. We find that for each vertex in the input graph, transformers compute the set of vertices reachable from that vertex. Each layer then progressively expands these sets, allowing the model to search over a number of vertices exponential in the number of layers. However, we find that as the input graph size increases, the transformer has greater difficulty in learning the task. This difficulty is not resolved even as the number of parameters is increased, suggesting that increasing model scale will not lead to robust search abilities. We also find that performing search in-context (i.e., chain-of-thought) does not resolve this inability to learn to search on larger graphs.
Abulhair Saparov, Srushti Pawar, Shreyas Pimpalgaonkar, Nitish Joshi, Richard Yuanzhe Pang, Vishakh Padmakumar, Mehran Kazemi, Najoung Kim, He He 0001
ICLR8
2025 Vision-and-Language Training Helps Deploy Taxonomic Knowledge but Does Not Fundamentally Alter It
abstract
Does vision-and-language (VL) training change the linguistic representations of language models in meaningful ways? In terms of downstream task performance on text-only tasks, most results in the literature have shown marginal differences. In this work, we start from the hypothesis that the domain in which VL training could have a significant effect is lexical-conceptual knowledge, in particular its taxonomic organization. Through comparing minimal pairs of text-only LMs and their VL-trained counterparts, we first show that the VL models often outperform their text-only counterparts on a text-only question-answering task that requires taxonomic understanding of concepts mentioned in the questions. Using an array of targeted behavioral and representational analyses, we show that the LMs and VLMs do not differ significantly in terms of their taxonomic knowledge itself, but they differ in how they represent questions that contain concepts in a taxonomic relation vs. a non-taxonomic relation. This implies that the taxonomic knowledge itself does not change substantially through additional VL training, but VL training does improve the deployment of this knowledge in the context of a specific task, even when the presentation of the task is purely linguistic.
Yulu Qin, Dheeraj Varghese, Adam Dahlgren Lindström, Lucia Donatelli, Kanishka Misra, Najoung Kim
NeurIPS6
2024 Beyond Thumbs Up/Down: Untangling Challenges of Fine-Grained Feedback for Text-to-Image Generation
abstract
Human feedback plays a critical role in learning and refining reward models for text-to-image generation, but the optimal form the feedback should take for learning an accurate reward function has not been conclusively established. This paper investigates the effectiveness of fine-grained feedback which captures nuanced distinctions in image quality and prompt-alignment, compared to traditional coarse-grained feedback (for example, thumbs up/down or ranking between a set of options). While fine-grained feedback holds promise, particularly for systems catering to diverse societal preferences, we show that demonstrating its superiority to coarse-grained feedback is not automatic. Through experiments on real and synthetic preference data, we surface the complexities of building effective models due to the interplay of model choice, feedback type, and the alignment between human judgment and computational interpretation. We identify key challenges in eliciting and utilizing fine-grained feedback, prompting a reassessment of its assumed benefits and practicality. Our findings -- e.g., that fine-grained feedback can lead to worse models for a fixed budget, in some settings; however, in controlled settings with known attributes, fine grained rewards can indeed be more helpful -- call for careful consideration of feedback attributes and potentially beckon novel modeling approaches to appropriately unlock the potential value of fine-grained feedback in-the-wild.
Katie Collins, Najoung Kim, Yonatan Bitton, Verena Rieser, Shayegan Omidshafiei, Yushi Hu, Sherol Chen, Senjuti Dutta, Minsuk Chang, Kimin Lee, Youwei Liang, Georgina Evans, Sahil Singla 0005, Gang Li 0021, Adrian Weller, Junfeng He, Deepak Ramachandran, Krishnamurthy Dvijotham
AIES (1)2
2024 Structural Generalization of Modification in Adult Learners of an Artificial Language
Najoung Kim, Paul Smolensky
CogSci1
2024 Personas as a Way to Model Truthfulness in Language Models
abstract
Large language models (LLMs) are trained on vast amounts of text from the internet, which contains both factual and misleading information about the world.While unintuitive from a classic view of language models, recent work has shown that the truth value of a statement can be elicited from the model's representations.This paper presents an explanation, persona hypothesis, for why LLMs appear to know the truth despite not being trained with truth labels.We hypothesize that the pretraining data is generated by groups of (un)truthful agents whose outputs share common features, and they form a (un)truthful persona.By training on this data, LMs can infer and represent the persona in its activation space.This allows the model to separate truth from falsehoods and controls the truthfulness of its generation.We show evidence for the persona hypothesis via two observations: (1) we can probe whether a model's answer will be truthful before it is generated; (2) finetuning a model on a set of true facts improves its truthfulness on unseen topics.Next, using arithmetics as a synthetic environment, we show that structures of the pretraining data are crucial for the model to infer the truthful persona.Overall, our findings suggest that models can exploit hierarchical structures in the data to learn abstract concepts like truthfulness.
Nitish Joshi, Javier Rando, Abulhair Saparov, Najoung Kim, He He 0001
EMNLP4
2024 Semantic Training Signals Promote Hierarchical Syntactic Generalization in Transformers
abstract
Neural networks without hierarchical biases often struggle to learn linguistic rules that come naturally to humans.However, neural networks are trained primarily on form alone, while children acquiring language additionally receive data about meaning.Would neural networks generalize more like humans when trained on both form and meaning?We investigate this by examining if Transformers-neural networks without a hierarchical bias-better achieve hierarchical generalization when trained on both form and meaning compared to when trained on form alone.Our results show that Transformers trained on form and meaning do favor the hierarchical generalization more than those trained on form alone, suggesting that statistical learners without hierarchical biases can leverage semantic training signals to bootstrap hierarchical syntactic generalization.
Aditya Yedetore, Najoung Kim
EMNLP2
2024 Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks
abstract
Zhaofeng Wu, Linlu Qiu, Alexis Ross, Ekin Akyürek, Boyuan Chen, Bailin Wang, Najoung Kim, Jacob Andreas, Yoon Kim. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Zhaofeng Wu, Linlu Qiu, Alexis Ross, Ekin Akyürek, Boyuan Chen 0003, Bailin Wang, Najoung Kim, Jacob Andreas
NAACL-HLT7
2023 LAMBADA: Backward Chaining for Automated Reasoning in Natural Language
abstract
Remarkable progress has been made on automated reasoning with natural text, by using Language Models (LMs) and methods such as Chain-of-Thought and Selection-Inference.These techniques search for proofs in the forward direction from axioms to the conclusion, which suffers from a combinatorial explosion of the search space, and thus high failure rates for problems requiring longer chains of reasoning.The classical automated reasoning literature has shown that reasoning in the backward direction (i.e. from the intended conclusion to supporting axioms) is significantly more efficient at proof-finding.Importing this intuition into the LM setting, we develop a Backward Chaining algorithm, called LAM-BADA, that decomposes reasoning into four sub-modules.These sub-modules are simply implemented by few-shot prompted LM inference.We show that LAMBADA achieves sizable accuracy boosts over state-of-the-art forward reasoning methods on two challenging logical reasoning datasets, particularly when deep and accurate proof chains are required. Facts:1. Rough and cold that is what they say about Blue Bob. 2. Eric, who is relatively young, is also pretty big and tends to be cold.3. Fred is green and cold too.4. For being so cold, it's good Harry can remain nice.Rules: 1. Rough, cold people are blue.2. Big, kind folks are green ones.3.If a person is big, rough, and cold, they are also red. 4. Most round and cold people are often rough.5. Cold, young people are also certain to be rough people.6.An individual who is big, red and young is also a nice individual.
Mehran Kazemi, Najoung Kim, Deepti Bhatia, Deepak Ramachandran
ACL (1)2
2023 (QA)²: Question Answering with Questionable Assumptions
abstract
Naturally occurring information-seeking questions often contain questionable assumptions-assumptions that are false or unverifiable.Questions containing questionable assumptions are challenging because they require a distinct answer strategy that deviates from typical answers for information-seeking questions.For instance, the question When did Marie Curie discover Uranium?cannot be answered as a typical when question without addressing the false assumption Marie Curie discovered Uranium.In this work, we propose (QA) 2 (Question Answering with Questionable Assumptions), an open-domain evaluation dataset consisting of naturally occurring search engine queries that may or may not contain questionable assumptions.To be successful on (QA) 2 , systems must be able to detect questionable assumptions and also be able to produce adequate responses for both typical information-seeking questions and ones with questionable assumptions.Through human rater acceptability on end-to-end QA with (QA) 2 , we find that current models do struggle with handling questionable assumptions, leaving substantial headroom for progress.* Equal contribution, corresponding authors ∆ Work partly done at NYU before joining BU. δ Work done at NYU before joining Amazon. 1 We use the term questionable assumptions instead of presupposition failure to capture failures of both true presup-
Najoung Kim, Phu Mon Htut, Samuel R. Bowman, Jackson Petty
ACL (1)1
2023 Entity Tracking in Language Models
abstract
Keeping track of how states of entities change as a text or dialog unfolds is a key prerequisite to discourse understanding.Yet, there have been few systematic investigations into the ability of large language models (LLMs) to track discourse entities.In this work, we present a task probing to what extent a language model can infer the final state of an entity given an English description of the initial state and a series of state-changing operations.We use this task to first investigate whether Flan-T5, GPT-3 and GPT-3.5 can track the state of entities, and find that only GPT-3.5 models, which have been pretrained on large amounts of code, exhibit this ability.We then investigate whether smaller models pretrained primarily on text can learn to track entities, through finetuning T5 on several training/evaluation splits.While performance degrades for more complex splits, we find that even when evaluated on a different set of entities from training or longer operation sequences, a finetuned model can perform nontrivial entity tracking.Taken together, these results suggest that language models can learn to track entities but pretraining on text corpora alone does not make this capacity surface.
Najoung Kim, Sebastian Schuster 0001
ACL (1)1
2023 SLOG: A Structural Generalization Benchmark for Semantic Parsing
abstract
The goal of compositional generalization benchmarks is to evaluate how well models generalize to new complex linguistic expressions.Existing benchmarks often focus on lexical generalization, the interpretation of novel lexical items in syntactic structures familiar from training.Structural generalization tasks, where a model needs to interpret syntactic structures that are themselves unfamiliar from training, are often underrepresented, resulting in overly optimistic perceptions of how well models can generalize.We introduce SLOG, a semantic parsing dataset that extends COGS (Kim and Linzen, 2020) with 17 structural generalization cases.In our experiments, the generalization accuracy of Transformer models, including pretrained ones, only reaches 40.6%, while a structure-aware parser only achieves 70.8%.These results are far from the near-perfect accuracy existing models achieve on COGS, demonstrating the role of SLOG in foregrounding the large discrepancy between models' lexical and structural generalization capacities.
Bingzhi Li, Lucia Donatelli, Alexander Koller, Tal Linzen, Yuekun Yao, Najoung Kim
EMNLP6
2023 Inverse Scaling Can Become U-Shaped
abstract
Scaling up language models has been empirically shown to improve performance on a wide range of downstream tasks.However, if we were to observe worse performance as a function of scale (inverse scaling) on certain tasks, this would indicate that scaling can also encourage behaviors that are misaligned with human preferences.The Inverse Scaling Prize (McKenzie et al., 2023) identified eleven such inverse scaling tasks, evaluated on models of up to 280B parameters and up to 500 zettaFLOPs of training compute.In this paper, we evaluate models of up to 540B parameters, trained on five times more compute than those evaluated in the Inverse Scaling Prize.With this increased range of model sizes and compute, only four out of the eleven tasks remain inverse scaling.Six tasks exhibit Ushaped scaling, where performance decreases up to a certain size, and then increases again up to the largest model evaluated (the one remaining task displays positive scaling).In addition, 1-shot examples and chain-of-thought can help mitigate undesirable scaling patterns even further.U-shaped scaling suggests that the inverse scaling trend observed in McKenzie et al. (2023) may not continue to hold for larger models, which we attribute to the presence of distractor tasks that only sufficiently large models can avoid.
Jason Wei, Najoung Kim, Yi Tay, Quoc V. Le
EMNLP2
2023 BoardgameQA: A Dataset for Natural Language Reasoning with Contradictory Information
abstract
Automated reasoning with unstructured natural text is a key requirement for many potential applications of NLP and for developing robust AI systems. Recently, Language Models (LMs) have demonstrated complex reasoning capacities even without any finetuning. However, existing evaluation for automated reasoning assumes access to a consistent and coherent set of information over which models reason. When reasoning in the real-world, the available information is frequently inconsistent or contradictory, and therefore models need to be equipped with a strategy to resolve such conflicts when they arise. One widely-applicable way of resolving conflicts is to impose preferences over information sources (e.g., based on source credibility or information recency) and adopt the source with higher preference. In this paper, we formulate the problem of reasoning with contradictory information guided by preferences over sources as the classical problem of defeasible reasoning, and develop a dataset called BoardgameQA for measuring the reasoning capacity of LMs in this setting. BoardgameQA also incorporates reasoning with implicit background knowledge, to better reflect reasoning problems in downstream applications. We benchmark various LMs on BoardgameQA and the results reveal a significant gap in the reasoning capacity of state-of-the-art LMs on this problem, showing that reasoning with conflicting information does not surface out-of-the-box in LMs. While performance can be improved with finetuning, it nevertheless remains poor.
Mehran Kazemi, Quan Yuan 0001, Deepti Bhatia, Najoung Kim, Vaiva Imbrasaite, Deepak Ramachandran
NeurIPS4
2023 Testing the General Deductive Reasoning Capacity of Large Language Models Using OOD Examples
abstract
Given the intractably large size of the space of proofs, any model that is capable of general deductive reasoning must generalize to proofs of greater complexity. Recent studies have shown that large language models (LLMs) possess some abstract deductive reasoning ability given chain-of-thought prompts. However, they have primarily been tested on proofs using modus ponens or of a specific size, and from the same distribution as the in-context examples. To measure the general deductive reasoning ability of LLMs, we test on a broad set of deduction rules and measure their ability to generalize to more complex proofs from simpler demonstrations from multiple angles: depth-, width-, and compositional generalization. To facilitate systematic exploration, we construct a new synthetic and programmable reasoning dataset that enables control over deduction rules and proof complexity. Our experiments on four LLMs of various sizes and training objectives show that they are able to generalize to compositional proofs. However, they have difficulty generalizing to longer proofs, and they require explicit demonstrations to produce hypothetical subproofs, specifically in proof by cases and proof by contradiction.
Abulhair Saparov, Richard Yuanzhe Pang, Vishakh Padmakumar, Nitish Joshi, Mehran Kazemi, Najoung Kim, He He 0001
NeurIPS6
2021 Which Linguist Invented the Lightbulb? Presupposition Verification for Question-Answering
abstract
Najoung Kim, Ellie Pavlick, Burcu Karagol Ayan, Deepak Ramachandran. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Najoung Kim, Ellie Pavlick, Burcu Karagol Ayan, Deepak Ramachandran
ACL/IJCNLP (1)1
2020 Implicit Discourse Relation Classification: We Need to Talk about Evaluation
abstract
Implicit relation classification onPenn Discourse TreeBank (PDTB) 2.0 is a common benchmark task for evaluating the understanding of discourse relations.However, the lack of consistency in preprocessing and evaluation poses challenges to fair comparison of results in the literature.In this work, we highlight these inconsistencies and propose an improved evaluation protocol.Paired with this protocol, we report strong baseline results from pretrained sentence encoders, which set the new state-of-the-art for PDTB 2.0.Furthermore, this work is the first to explore fine-grained relation classification on PDTB 3.0.We expect our work to serve as a point of comparison for future work, and also as an initiative to discuss models of larger context and possible data augmentations for downstream transferability.
Najoung Kim, Song Feng 0002, R. Chulaka Gunasekara, Luis A. Lastras
ACL1
2020 COGS: A Compositional Generalization Challenge Based on Semantic Interpretation
abstract
Natural language is characterized by compositionality: the meaning of a complex expression is constructed from the meanings of its constituent parts.To facilitate the evaluation of the compositional abilities of language processing architectures, we introduce COGS, a semantic parsing dataset based on a fragment of English.The evaluation portion of COGS contains multiple systematic gaps that can only be addressed by compositional generalization; these include new combinations of familiar syntactic structures, or new combinations of familiar words and familiar structures.In experiments with Transformers and LSTMs, we found that in-distribution accuracy on the COGS test set was near-perfect (96-99%), but generalization accuracy was substantially lower (16-35%) and showed high sensitivity to random seed (±6-8%).These findings indicate that contemporary standard NLP models are limited in their compositional generalization capacity, and position COGS as a good way to measure progress.
Najoung Kim, Tal Linzen
EMNLP (1)1
2019 Predicting the Argumenthood of English Prepositional Phrases
Najoung Kim, Kyle Rawlins, Benjamin Van Durme, Paul Smolensky
AAAI1
2019 Can You Tell Me How to Get Past Sesame Street? Sentence-Level Pretraining Beyond Language Modeling
abstract
Alex Wang, Jan Hula, Patrick Xia, Raghavendra Pappagari, R. Thomas McCoy, Roma Patel, Najoung Kim, Ian Tenney, Yinghui Huang, Katherin Yu, Shuning Jin, Berlin Chen, Benjamin Van Durme, Edouard Grave, Ellie Pavlick, Samuel R. Bowman. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Jan Hula, Patrick Xia 0002, Raghavendra Pappagari, Tom McCoy 0001, Roma Patel, Najoung Kim, Ian Tenney, Katherin Yu, Shuning Jin, Berlin Chen, Benjamin Van Durme, Edouard Grave, Ellie Pavlick, Samuel R. Bowman
ACL (1)7
2019 What do you learn from context? Probing for sentence structure in contextualized word representations
Ian Tenney, Patrick Xia 0002, Berlin Chen, Adam Poliak, Tom McCoy 0001, Najoung Kim, Benjamin Van Durme, Samuel R. Bowman, Dipanjan Das 0001, Ellie Pavlick
ICLR (Poster)7
2016 Prosodic and Linguistic Analysis of Semantic Fluency Data: A Window into Speech Production and Cognition
abstract
Semantic fluency is a commonly used task in psychology that provides data about executive function and semantic memory. Performance on the task is affected by conditions ranging from depression to dementia. The task involves participants naming as many members of a given category (e.g. animals) as possible in sixty seconds. Most of the analyses reported in the literature only rely on word counts and transcribed data, and do not take into account the evidence of utterance planning present in the speech signal. Using data from Korean, we show how prosodic analyses can be combined with computational linguistic analyses of the words produced to provide further insights into the processes involved in producing fluency data. We compare our analyses to an established analysis method for semantic fluency data, manual determination of lexically coherent clusters of words.
Maria Klara Wolters, Najoung Kim, Jung-Ho Kim 0002, Sarah E. MacPherson, Jong-Chan Park
INTERSPEECH2