Tianyu Jiang 0001

dblp:46/11004-1 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
12since 2021 · last 2026
0009-0003-6019-2075ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 4 first-author · 12 since 2021
YearPublicationVenuePosition
2026 MetFuse: Figurative Fusion between Metonymy and Metaphor
abstract
Metonymy and metaphor often co-occur in natural language, yet computational work has studied them largely in isolation.We introduce a framework that transforms a literal sentence into three figurative variants: metonymic, metaphoric, and hybrid.Using this framework, we construct MetFuse, 1 the first dedicated dataset of figurative fusion between metonymy and metaphor, containing 1,000 human-verified meaning-aligned quadruplets totaling 4,000 sentences.Extrinsic experiments on eight existing benchmarks show that augmenting training data with MetFuse consistently improves both metonymy and metaphor classification, with hybrid examples yielding the largest gains on metonymy tasks.Using this dataset, we also analyze how the presence of one figurative type impacts another.Our findings show that both human annotators and large language models better identify metonymy in hybrid sentences than in metonymy-only sentences, demonstrating that the presence of a metaphor makes a metonymic noun more explicit.
Tianyu Jiang 0001
ACL (1)2
2026 Exploring Concreteness Through a Figurative Lens
abstract
Static concreteness ratings are widely used in NLP, yet a word's concreteness can shift with context, especially in figurative language such as metaphor, where common concrete nouns can take abstract interpretations.While such shifts are evident from context, it remains unclear how LLMs understand concreteness internally.We conduct a layer-wise and geometric analysis of LLM hidden representations across four model families, examining how models distinguish literal vs. figurative uses of the same noun and how concreteness is organized in representation space.We find that LLMs separate literal and figurative usage in early layers, and that mid-to-late layers compress concreteness into a one-dimensional direction that is consistent across models.Finally, we show this geometric structure is practically useful: a single concreteness direction supports efficient figurative-language classification and enables training-free steering of generation toward more literal or more figurative rewrites.
Tianyu Jiang 0001
ACL (1)2
2026 Evaluating the Impact of Verbal Multiword Expressions on Machine Translation
abstract
Verbal multiword expressions (VMWEs) remain difficult for machine translation because their meanings are often not recoverable from their component words.In this study, we analyze the impact of three VMWE categoriesverbal idioms, verb-particle constructions, and light verb constructions-on machine translation quality from English to multiple languages.Using both established multiword expression datasets and standard machine translation datasets, we evaluate how state-of-the-art translation systems handle these expressions.Our experimental results consistently show that VMWEs negatively affect translation quality, with deeper analysis indicating that this degradation is primarily attributable to the VMWE itself rather than general sentence-level difficulty.We release our code and evaluation framework to test new MT systems for the community.
Linfeng Liu 0006, Tianyu Jiang 0001
ACL (1)3
2026 Rhetorical Questions in LLM Representations: A Linear Probing Study
abstract
Rhetorical questions are asked not to seek information but to persuade or signal stance.How large language models internally represent them remains unclear.We analyze rhetorical questions in LLM representations using linear probes on two social-media datasets with different discourse contexts, and find that rhetorical signals emerge early and are most stably captured by last-token representations.Rhetorical questions are linearly separable from information-seeking questions within datasets, and remain detectable under crossdataset transfer, reaching AUROC around 0.7-0.8.However, we demonstrate that transferability does not simply imply a shared representation.Probes trained on different datasets produce different rankings when applied to the same target corpus, with overlap among the topranked instances often below 0.2.Qualitative analysis shows that these divergences correspond to distinct rhetorical phenomena: some probes capture discourse-level rhetorical stance embedded in extended argumentation, while others emphasize localized, syntax-driven interrogative acts.Together, these findings suggest that rhetorical questions in LLM representations are encoded by multiple linear directions emphasizing different cues, rather than a single shared direction.
Louie Hong Yao, Vishesh Anand, Tianyu Jiang 0001
ACL (1)4
2026 Beyond Supervised Clarification: Input Rewriting with LLMs for Dialogue Discourse Parsing
abstract
Rewriting inputs to improve frozen downstream models has become a common strategy in modern NLP pipelines. Prior work on incremental dialogue discourse parsing (DDP) shows that supervised clarification models can rewrite fragmentary or underspecified utterances—such as resolving ellipsis or references—to improve parsing accuracy. In this work, we revisit this idea under realistic deployment conditions, where no clarification supervision is available and the clarifier must rely on zero-shot prompting or feedback from a frozen parser. Across three Segmented Discourse Representation Theory (SDRT) datasets and multiple parsers, we find that last-utterance clarification is far less reliable than suggested by supervised settings. Parser-agnostic rewriting often introduces more regressions than repairs, as edits that enable fixes also disrupt discourse cues relied upon by the parser. A best-of-8 rewriting analysis further reveals a practical ceiling: a large fraction of errors are not repairable through input rewriting alone. A parser-aware clarifier trained with GRPO reduces regressions by up to 37% by learning conservative abstention, yet still fails to produce selectivity-aware clarifications that consistently improve parsing. Together, these findings recast clarification as a selective intervention problem. We identify rewritability prediction—deciding whether an utterance is repairable before intervention—as the key missing capability for input-side optimization of frozen discourse parsers, and a critical direction for improving agentic pipelines more broadly.[Data and code are available at https://github.com/ounlp/Clarification-for-DDP.]
Zhichao Xu 0001, Xin Yu 0002, Yingheng Tang, Tianyu Jiang 0001, Jie Cao 0010
SIGDIAL6
2025 Do LLMs Encode Frame Semantics? Evidence from Frame Identification
abstract
We investigate whether large language models encode latent knowledge of frame semantics, focusing on frame identification, a core challenge in frame semantic parsing that involves selecting the appropriate semantic frame for a target word in context.Using the FrameNet lexical resource, we evaluate models under prompt-based inference and observe that they can perform frame identification effectively even without explicit supervision.To assess the impact of task-specific training, we fine-tune the model on FrameNet data, which substantially improves in-domain accuracy while generalizing well to out-of-domain benchmarks.Further analysis shows that the models can generate semantically coherent frame definitions, highlighting the model's internalized understanding of frame semantics.
Jayanth Krishna Chundru, Rudrashis Poddar, Jie Cao 0010, Tianyu Jiang 0001
EMNLP4
2025 GuessingGame: Measuring the Informativeness of Open-Ended Questions in Large Language Models
abstract
We introduce GuessingGame, a protocol for evaluating large language models (LLMs) as strategic question-askers in open-ended, opendomain settings.A Guesser LLM identifies a hidden object by posing free-form questions to an Oracle without predefined choices or candidate lists.To measure question quality, we propose two information gain (IG) metrics: a Bayesian method that tracks belief updates over semantic concepts using LLM-scored relevance, and an entropy-based method that filters candidates via ConceptNet.Both metrics are model-agnostic and support post hoc analysis.Across 858 games with multiple models and prompting strategies, higher IG strongly predicts efficiency: a one-standard-deviation IG increase reduces expected game length by 43%.Prompting constraints guided by IG, such as enforcing question diversity, enable weaker models to significantly improve performance.These results show that question-asking in LLMs is both measurable and improvable, and crucial for interactive reasoning.
Dylan Hutson, Daniel Vennemeyer, Aneesh Deshmukh, Justin Zhijun Zhan, Tianyu Jiang 0001
EMNLP5
2025 ConMeC: A Dataset for Metonymy Resolution with Common Nouns
abstract
Saptarshi Ghosh, Tianyu Jiang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Tianyu Jiang 0001
NAACL (Long Papers)2
2024 My Heart Skipped a Beat! Recognizing Expressions of Embodied Emotion in Natural Language
abstract
Yuan Zhuang, Tianyu Jiang, Ellen Riloff. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Tianyu Jiang 0001, Ellen Riloff
NAACL-HLT2
2022 Identifying Physical Object Use in Sentences
abstract
Commonsense knowledge about the typical functions of physical objects allows people to make inferences during sentence understanding.For example, we infer that "Sam enjoyed the book" means that Sam enjoyed reading the book, even though the action is implicit.Prior research has focused on learning the prototypical functions of physical objects in order to enable inferences about implicit actions.But many sentences refer to objects even when they are not used (e.g., "The book fell").We argue that NLP systems need to recognize whether an object is being used before inferring how the object is used.We define a new task called Object Use Classification that determines whether a physical object mentioned in a sentence was used or likely will be used.We introduce a new dataset for this task and present a classification model that exploits data augmentation methods and FrameNet when fine-tuning a pre-trained language model.We also show that object use classification combined with knowledge about the prototypical functions of objects has the potential to yield very good inferences about implicit and anticipated actions.
Tianyu Jiang 0001, Ellen Riloff
EMNLP1
2021 Learning Prototypical Functions for Physical Artifacts
abstract
Tianyu Jiang, Ellen Riloff. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Tianyu Jiang 0001, Ellen Riloff
ACL/IJCNLP (1)1
2021 Exploiting Definitions for Frame Identification
abstract
Frame identification is one of the key challenges for frame-semantic parsing.The goal of this task is to determine which frame best captures the meaning of a target word or phrase in a sentence.We present a new model for frame identification that uses a pre-trained transformer model to generate representations for frames and lexical units (senses) using their formal definitions in FrameNet.Our frame identification model assesses the suitability of a frame for a target word in a sentence based on the semantic coherence of their meanings.We evaluate our model on three data sets and show that it consistently achieves better performance than previous systems.
Tianyu Jiang 0001, Ellen Riloff
EACL1
2020 Affective Event Classification with Discourse-enhanced Self-training
abstract
Prior research has recognized the need to associate affective polarities with events and has produced several techniques and lexical resources for identifying affective events.Our research introduces new classification models to assign affective polarity to event phrases.First, we present a BERT-based model for affective event classification and show that the classifier achieves substantially better performance than a large affective event knowledge base.Second, we present a discourse-enhanced selftraining method that iteratively improves the classifier with unlabeled data.The key idea is to exploit event phrases that occur with a coreferent sentiment expression.The discourseenhanced self-training algorithm iteratively labels new event phrases based on both the classifier's predictions and the polarities of the event's coreferent sentiment expressions.Our results show that discourse-enhanced selftraining further improves both recall and precision for affective event classification.
Tianyu Jiang 0001, Ellen Riloff
EMNLP (1)2
2018 Learning Prototypical Goal Activities for Locations
abstract
People go to different places to engage in activities that reflect their goals.For example, people go to restaurants to eat, libraries to study, and churches to pray.We refer to an activity that represents a common reason why people typically go to a location as a prototypical goal activity (goal-act).Our research aims to learn goal-acts for specific locations using a text corpus and semi-supervised learning.First, we extract activities and locations that co-occur in goal-oriented syntactic patterns.Next, we create an activity profile matrix and apply a semi-supervised label propagation algorithm to iteratively revise the activity strengths for different locations using a small set of labeled data.We show that this approach outperforms several baseline methods when judged against goal-acts identified by human annotators.
Tianyu Jiang 0001, Ellen Riloff
ACL (1)1