Jingxuan Tu

dblp:120/2772 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author
YearPublicationVenuePosition
2026 Not All Disneys Are the Same: Making Coreference Metonymy-Aware
Bingyang Ye, Jingxuan Tu, James Pustejovsky
LREC2
2025 Enhanced Noun-Noun Compound Interpretation through Textual Enrichment
abstract
Interpreting Noun-Noun Compounds remains a persistent challenge for Large Language Models (LLMs) because the semantic relation between the modifier and the head is rarely stated explicitly.Recent benchmarks frame Noun-Noun Compound Interpretation as a multiplechoice question.While this setting allows LLMs to produce more controlled results, it still faces two key limitations: vague relation descriptions as options and the inability to handle polysemous compounds.We introduce a dual-faceted textual enrichment framework that augments prompts.Description enrichment paraphrases relations into event-oriented descriptions instantiated with the target compound to explicitly surface the hidden event connecting head and modifier.Conditioned context enrichment identifies polysemous compounds leveraging qualia-role binding and assigns each compound with condition cues for disambiguation.Our method yields consistently higher accuracy across three LLM families.These gains suggest that surfacing latent compositional structure and contextual constraint is a promising path toward deeper semantic understanding in language models. 1
Bingyang Ye, Jingxuan Tu, James Pustejovsky
EMNLP2
2025 Beyond Benchmarks: Building a Richer Cross-Document Event Coreference Dataset with Decontextualization
abstract
Jin Zhao, Jingxuan Tu, Bingyang Ye, Xinrui Hu, Nianwen Xue, James Pustejovsky. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Jin Zhao 0009, Jingxuan Tu, Bingyang Ye, Xinrui Hu, Nianwen Xue, James Pustejovsky
NAACL (Long Papers)2
2024 Linguistically Conditioned Semantic Textual Similarity
abstract
Semantic textual similarity (STS) is a fundamental NLP task that measures the semantic similarity between a pair of sentences.In order to reduce the inherent ambiguity posed from the sentences, a recent work called Conditional STS (C-STS) has been proposed to measure the sentences' similarity conditioned on a certain aspect.Despite the popularity of C-STS, we find that the current C-STS dataset suffers from various issues that could impede proper evaluation on this task.In this paper, we reannotate the C-STS validation set and observe an annotator discrepancy on 55% of the instances resulting from the annotation errors in the original label, ill-defined conditions, and the lack of clarity in the task definition.After a thorough dataset analysis, we improve the C-STS task by leveraging the models' capability to understand the conditions under a QA task setting.With the generated answers, we present an automatic error identification pipeline that is able to identify annotation errors from the C-STS data with over 80% F1 score.We also propose a new method that largely improves the performance over baselines on the C-STS data by training the models with the answers.Finally we discuss the conditionality annotation based on the typed-feature structure (TFS) of entity types.We show in examples that the TFS is able to provide a linguistic foundation for constructing C-STS data with new conditions.
Jingxuan Tu, Keer Xu, Liulu Yue, Bingyang Ye, Kyeongmin Rim, James Pustejovsky
ACL (1)1
2024 Common Ground Tracking in Multimodal Dialogue
abstract
Within Dialogue Modeling research in AI and NLP, considerable attention has been spent on “dialogue state tracking” (DST), which is the ability to update the representations of the speaker’s needs at each turn in the dialogue by taking into account the past dialogue moves and history. Less studied but just as important to dialogue modeling, however, is “common ground tracking” (CGT), which identifies the shared belief space held by all of the participants in a task-oriented dialogue: the task-relevant propositions all participants accept as true. In this paper we present a method for automatically identifying the current set of shared beliefs and ”questions under discussion” (QUDs) of a group with a shared goal. We annotate a dataset of multimodal interactions in a shared physical space with speech transcriptions, prosodic features, gestures, actions, and facets of collaboration, and operationalize these features for use in a deep neural model to predict moves toward construction of common ground. Model outputs cascade into a set of formal closure rules derived from situated evidence and belief axioms and update operations. We empirically assess the contribution of each feature type toward successful construction of common ground relative to ground truth, establishing a benchmark in this novel, challenging task.
Ibrahim Khebour, Kenneth Lai, Mariah Bradford, Yifan Zhu 0014, Richard Brutti, Christopher Tam, Jingxuan Tu, Benjamin Ibarra, Nathaniel Blanchard, Nikhil Krishnaswamy, James Pustejovsky
LREC/COLING7
2024 GLAMR: Augmenting AMR with GL-VerbNet Event Structure
abstract
This paper introduces GLAMR, an Abstract Meaning Representation (AMR) interpretation of Generative Lexicon (GL) semantic components. It includes a structured subeventual interpretation of linguistic predicates, and encoding of the opposition structure of property changes of event arguments. Both of these features are recently encoded in VerbNet (VN), and form the scaffolding for the semantic form associated with VN frame files. We develop a new syntax, concepts, and roles for subevent structure based on VN for connecting subevents to atomic predicates. Our proposed extension is compatible with current AMR specification. We also present an approach to automatically augment AMR graphs by inserting subevent structure of the predicates and identifying the subevent arguments from the semantic roles. A pilot annotation of GLAMR graphs of 65 documents (486 sentences), based on procedural texts as a source, is presented as a public dataset. The annotation includes subevents, argument property change, and document-level anaphoric links. Finally, we provide baseline models for converting text to GLAMR and vice versa, along with the application of GLAMR for generating enriched paraphrases with details on subevent transformation and arguments that are not present in the surface form of the texts.
Jingxuan Tu, Timothy Obiso, Bingyang Ye, Kyeongmin Rim, Keer Xu, Liulu Yue, Susan Windisch Brown, Martha Palmer, James Pustejovsky
LREC/COLING1
2024 Propositional Extraction from Natural Speech in Small Group Collaborative Tasks
Videep Venkatesha, Abhijnan Nath, Ibrahim Khebour, Avyakta Chelle, Mariah Bradford, Jingxuan Tu, James Pustejovsky, Nathaniel Blanchard, Nikhil Krishnaswamy
EDM6
2024 Media Attitude Detection via Framing Analysis with Events and their Relations
abstract
Framing is used to present some selective aspects of an issue and make them more salient, which aims to promote certain values, interpretations, or solutions (Entman, 1993).This study investigates the nuances of media framing on public perception and understanding by examining how events are presented within news articles.Unlike previous research that primarily focused on word choice as a framing device, this work explores the comprehensive narrative construction through events and their relations.Our method integrates event extraction, Cross-Document Event Coreference (CDEC), and causal relationship among events to extract framing devices employed by the media to assess their role in framing the narrative.We evaluate our approach with a media attitude detection task and show that the use of event mentions, event cluster descriptors, and their causal relations effectively captures the subtle nuances of framing, thereby providing deeper insights into the attitudes conveyed by news articles.The experimental results show the framing device models surpass the baseline models and offer a more detailed and explainable analysis of media framing effects.We make the source code and dataset publicly available.1
Jin Zhao 0009, Jingxuan Tu, Nianwen Xue
EMNLP2
2022 Interpreting Logical Metonymy through Dense Paraphrasing
Bingyang Ye, Jingxuan Tu, Elisabetta Jezek, James Pustejovsky
CogSci2
2022 Competence-based Question Generation
abstract
Models of natural language understanding often rely on question answering and logical inference benchmark challenges to evaluate the performance of a system. While informative, such task-oriented evaluations do not assess the broader semantic abilities that humans have as part of their linguistic competence when speaking and interpreting language. We define competence-based (CB) question generation, and focus on queries over lexical semantic knowledge involving implicit argument and subevent structure of verbs. We present a method to generate such questions and a dataset of English cooking recipes we use for implementing the generation method. Our primary experiment shows that even large pretrained language models perform poorly on CB questions until they are provided with additional contextualized semantic information. The data and the source code is available at: https://github.com/brandeis-llc/CompQG.
Jingxuan Tu, Kyeongmin Rim, James Pustejovsky
COLING1
2022 Evaluating Retrieval for Multi-domain Scientific Publications
abstract
This paper provides an overview of the xDD/LAPPS Grid framework and provides results of evaluating the AskMe retrievalengine using the BEIR benchmark datasets. Our primary goal is to determine a solid baseline of performance to guide furtherdevelopment of our retrieval capabilities. Beyond this, we aim to dig deeper to determine when and why certain approachesperform well (or badly) on both in-domain and out-of-domain data, an issue that has to date received relatively little attention.
Nancy Ide, Keith Suderman, Jingxuan Tu, Marc Verhagen, Shanan Peters, John Lawson, Andrew Borg, James Pustejovsky
LREC3
2020 Reproducing Neural Ensemble Classifier for Semantic Relation Extraction inScientific Papers
abstract
Within the natural language processing (NLP) community, shared tasks play an important role. They define a common goal and allowthe the comparison of different methods on the same data. SemEval-2018 Task 7 involves the identification and classification of relationsin abstracts from computational linguistics (CL) publications. In this paper we describe an attempt to reproduce the methods and resultsfrom the top performing system at for SemEval-2018 Task 7. We describe challenges we encountered in the process, report on the resultsof our system, and discuss the ways that our attempt at reproduction can inform best practices.
Kyeongmin Rim, Jingxuan Tu, Kelley Lynch, James Pustejovsky
LREC2
2019 On the analysis of spectrum based fault localization using hitting sets
Jingxuan Tu, Xiaoyuan Xie, Tsong Yueh Chen, Baowen Xu
J. Syst. Softw.1
2016 Code Coverage-Based Failure Proximity without Test Oracles
abstract
Failure indexing technique plays an important role in modern software maintenance. It can facilitate duplicated failure removal, failure assignment, etc. Failure proximity is a crucial part that underpins failure indexing techniques. It is comprised of two components: a fingerprinting function extracting failure signatures from failures and a distance function computing pairwise distances between failures. Failure proximity usually assumes the existence of test oracle. However, in many real-life application domains, test oracles do not always exist. Hence, the applicability of existing failure proximity techniques is limited. In our paper, we focus on investigating how to apply metamorphic testing on code coverage-based failure proximity without test oracles. In our approach, instead of using the testing results of failure, the results of violation or non-violation for metamorphic test groups are used. Specifically, the fingerprinting function extracts signatures from metamorphic slices rather than execution slices and the distance function computes the pairwise distance between violations rather than between failures. Thereby, the applicability of failure proximity is extended to the situations without test oracles. The experimental results on 50 two-fault mutants show that the quality of proximity matrix obtained through our approach is statistical comparable to traditional code coverage-based failure proximity with test oracle.
Jingxuan Tu, Xiaoyuan Xie, Baowen Xu
COMPSAC1