Xiangjue Dong

dblp:266/1362 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0002-9173-2690ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CHOIR: Harmonizing Structured Persona Diversity for Robust Collaborative LLM Reasoning
abstract
Xiangjue Dong, Cong Wang, Maria Teleki, Millennium Bismay, Ruihong Huang, James Caverlee. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xiangjue Dong, Maria Teleki, Millennium Bismay, Ruihong Huang, James Caverlee
ACL (1)1
2026 DMRetriever: A Family of Models for Improved Text Retrieval in Disaster Management
abstract
Kai Yin, Xiangjue Dong, Chengkai Liu, Allen Lin, Lingfeng Shi, Ali Mostafavi, James Caverlee. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xiangjue Dong, Chengkai Liu, Allen Lin, Lingfeng Shi, Ali Mostafavi, James Caverlee
ACL (1)2
2026 Multi-Scale Model Compression via Nested Matrix Learning
Xiangjue Dong, Aditya Anantharaman, Hemant Pugaliya
LREC1
2026 Language Models as Semantic Augmenters for Sequential Recommenders
abstract
Large Language Models (LLMs) excel at capturing latent semantics and contextual relationships across diverse modalities. However, in modeling user behavior from sequential interaction data, performance often suffers when such semantic context is limited or absent. We introduce LaMAR, a LLM-driven semantic enrichment framework designed to enrich such sequences automatically. LaMAR leverages LLMs in a few-shot setting to generate auxiliary contextual signals by inferring latent semantic aspects of a user's intent and item relationships from existing metadata. These generated signals, such as inferred usage scenarios, item intents, or thematic summaries, augment the original sequences with greater contextual depth. We demonstrate the utility of this generated resource by integrating it into benchmark sequential modeling tasks, where it consistently improves performance. Further analysis shows that LLM-generated signals exhibit high semantic novelty and diversity, enhancing the representational capacity of the downstream models. This work represents a new data-centric paradigm where LLMs serve as intelligent context generators, contributing a new method for the semi-automatic creation of training data and language resources.
Mahsa Valizadeh, Xiangjue Dong, Rui Tuo 0001, James Caverlee
LREC2
2025 Masculine Defaults via Gendered Discourse in Podcasts and Large Language Models
abstract
Masculine defaults are widely recognized as a significant type of gender bias, but they are often unseen as they are under-researched. Masculine defaults involve three key parts: (i) the cultural context, (ii) the masculine characteristics or behaviors, and (iii) the reward for, or simply acceptance of, those masculine characteristics or behaviors. In this work, we study discourse-based masculine defaults, and propose a twofold framework for (i) the large-scale discovery and analysis of gendered discourse words in spoken content via our Gendered Discourse Correlation Framework (GDCF); and (ii) the measurement of the gender bias associated with these gendered discourse words in LLMs via our Discourse Word-Embedding Association Test (D-WEAT). We focus our study on podcasts, a popular and growing form of social media, analyzing 15,117 podcast episodes. We analyze correlations between gender and discourse words -- discovered via LDA and BERTopic -- to automatically form gendered discourse word lists. We then study the prevalence of these gendered discourse words in domain-specific contexts, and find that gendered discourse-based masculine defaults exist in the domains of business, technology/politics, and video games. Next, we study the representation of these gendered discourse words from a state-of-the-art LLM embedding model from OpenAI, and find that the masculine discourse words have a more stable and robust representation than the feminine discourse words, which may result in better system performance on downstream tasks for men. Hence, men are rewarded for their discourse patterns with better system performance by one of the state-of-the-art language models -- and this embedding disparity is a representational harm and a masculine default.
Maria Teleki, Xiangjue Dong, James Caverlee
ICWSM2
2024 DACL: Disfluency Augmented Curriculum Learning for Fluent Text Generation
abstract
Voice-driven software systems are in abundance. However, language models that power these systems are traditionally trained on fluent, written text corpora. Hence there can be a misalignment between the inherent disfluency of transcribed spoken content and the fluency of the written training data. Furthermore, gold-standard disfluency annotations of various complexities for incremental training can be expensive to collect. So, we propose in this paper a Disfluency Augmented Curriculum Learning (DACL) approach to tackle the complex structure of disfluent sentences and generate fluent texts from them, by using Curriculum Learning (CL) coupled with our synthetically augmented disfluent texts of various levels. DACL harnesses the tiered structure of our generated synthetic disfluent data using CL, by training the model on basic samples (i.e. more fluent) first before training it on more complex samples (i.e. more disfluent). In contrast to the random data exposure paradigm, DACL focuses on a simple-to-complex learning process. We comprehensively evaluate DACL on Switchboard Penn Treebank-3 and compare it to the state-of-the-art disfluency removal models. Our model surpasses existing techniques in word-based precision (by up to 1%) and has shown favorable recall and F1 scores.
Rohan Chaudhury, Maria Teleki, Xiangjue Dong, James Caverlee
LREC/COLING3
2024 Quantifying the Impact of Disfluency on Spoken Content Summarization
abstract
Spoken content is abundant – including podcasts, meeting transcripts, and TikTok-like short videos. And yet, many important tasks like summarization are often designed for written content rather than the looser, noiser, and more disfluent style of spoken content. Hence, we aim in this paper to quantify the impact of disfluency on spoken content summarization. Do disfluencies negatively impact the quality of summaries generated by existing approaches? And if so, to what degree? Coupled with these goals, we also investigate two methods towards improving summarization in the presence of such disfluencies. We find that summarization quality does degrade with an increase in these disfluencies and that a combination of multiple disfluency types leads to even greater degradation. Further, our experimental results show that naively removing disfluencies and augmenting with special tags can worsen the summarization when used for testing, but that removing disfluencies for fine-tuning yields the best results. We make the code available at https://github.com/mariateleki/Quantifying-Impact-Disfluency.
Maria Teleki, Xiangjue Dong, James Caverlee
LREC/COLING2
2024 The Neglected Tails in Vision-Language Models
abstract
Vision-language models (VLMs) excel in zero-shot recognition but their performance varies greatly across different visual concepts. For example, although CLIP achieves impressive accuracy on ImageNet (60-80%), its performance drops below 10% for more than ten concepts like night snake, presumably due to their limited presence in the pretraining data. However, measuring the frequency of concepts in VLMs' large-scale datasets is challenging. We address this by using large language models (LLMs) to count the number of pretraining texts that con-tain synonyms of these concepts. Our analysis confirms that popular datasets, such as LAION, exhibit a long-tailed concept distribution, yielding biased performance in VLMs. We also find that downstream applications of VLMs, including visual chatbots (e.g., GPT-4V) and text-to-image models (e.g., Stable Diffusion), often fail to recognize or generate images of rare concepts identified by our method. To mit-igate the imbalanced performance of zero-shot VLMs, we propose REtrieval-Augmented Learning (REAL). First, in-stead of prompting VLMs using the original class names, REAL uses their most frequent synonyms found in pretraining texts. This simple change already outperforms costly human-engineered and LLM-enriched prompts over nine benchmark datasets. Second, REAL trains a linear classifier on a small yet balanced set of pretraining data re-trieved using concept synonyms. REAL surpasses the previous zero-shot SOTA, using 400× less storage and 10,000× less training time!
Shubham Parashar, Zhiqiu Lin, Tian Liu 0006, Xiangjue Dong, Deva Ramanan, James Caverlee, Shu Kong
CVPR4
2024 DA³: A Distribution-Aware Adversarial Attack against Language Models
abstract
Language models can be manipulated by adversarial attacks, which introduce subtle perturbations to input data.While recent attack methods can achieve a relatively high attack success rate (ASR), we've observed that the generated adversarial examples have a different data distribution compared with the original examples.Specifically, these adversarial examples exhibit reduced confidence levels and greater divergence from the training data distribution.Consequently, they are easy to detect using straightforward detection methods, diminishing the efficacy of such attacks.To address this issue, we propose a Distribution-Aware Adversarial Attack (DA 3 ) method.DA 3 considers the distribution shifts of adversarial examples to improve attacks' effectiveness under detection methods.We further design a novel evaluation metric, the Non-detectable Attack Success Rate (NASR), which integrates both ASR and detectability for the attack task.We conduct experiments on four widely used datasets to validate the attack effectiveness and transferability of adversarial examples generated by DA 3 against both the white-box BERT-BASE and ROBERTA-BASE models and the black-box LLAMA2-7B model 1 .
Yibo Wang 0001, Xiangjue Dong, James Caverlee, Philip S. Yu
EMNLP2
2024 Comparing ASR Systems in the Context of Speech Disfluencies
Maria Teleki, Xiangjue Dong, Soohwan Kim, James Caverlee
INTERSPEECH2
2023 Closed-book Question Generation via Contrastive Learning
abstract
Question Generation (QG) is a fundamental NLP task for many downstream applications.Recent studies on open-book QG, where supportive answer-context pairs are provided to models, have achieved promising progress.However, generating natural questions under a more practical closed-book setting that lacks these supporting documents still remains a challenge.In this work, we propose a new QG model for this closed-book setting that is designed to better understand the semantics of long-form abstractive answers and store more information in its parameters through contrastive learning and an answer reconstruction module.Through experiments, we validate the proposed QG model on both public datasets and a new WikiCQA dataset.Empirical results show that the proposed QG model outperforms baselines in both automatic evaluation and human evaluation.In addition, we show how to leverage the proposed model to improve existing question-answering systems.These results further indicate the effectiveness of our QG model for enhancing closed-book questionanswering tasks.
Xiangjue Dong, Jiaying Lu 0001, Jianling Wang, James Caverlee
EACL1
2023 Weakly Supervised Concept Map Generation Through Task-Guided Graph Translation
abstract
Recent years have witnessed the rapid development of concept map generation techniques due to their advantages in providing well-structured summarization of knowledge from free texts. Traditional unsupervised methods do not generate task-oriented concept maps, whereas deep generative models require large amounts of training data. In this work, we presentGT-D2G(Graph Translation-based Document To Graph), an automatic concept map generation framework that leverages generalized NLP pipelines to derive semantic-rich initial graphs, and translates them into more concise structures under the weak supervision of downstream task labels. The concept maps generated byGT-D2Gcan provide interpretable summarization of structured knowledge for the input texts, which are demonstrated through human evaluation and case studies on three real-world corpora. Further experiments on the downstream task of document classification show thatGT-D2Gbeats other concept map generation methods. Moreover, we specifically validate the labeling efficiency ofGT-D2Gin the label-efficient learning setting and the flexibility of generated graph sizes in controlled hyper-parameter studies.
Jiaying Lu 0001, Xiangjue Dong, Carl Yang 0001
IEEE Trans. Knowl. Data Eng.2