VLDB 2026 Research / reviewers in the wild / expert
Xiangjue Dong
dblp:266/1362
· DBLP profile ↗
12ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0002-9173-2690ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CHOIR: Harmonizing Structured Persona Diversity for Robust Collaborative LLM ReasoningabstractXiangjue Dong, Cong Wang, Maria Teleki, Millennium Bismay, Ruihong Huang, James Caverlee. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xiangjue Dong, Maria Teleki, Millennium Bismay, Ruihong Huang, James Caverlee |
ACL (1) | 1 |
| 2026 | DMRetriever: A Family of Models for Improved Text Retrieval in Disaster ManagementabstractKai Yin, Xiangjue Dong, Chengkai Liu, Allen Lin, Lingfeng Shi, Ali Mostafavi, James Caverlee. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xiangjue Dong, Chengkai Liu, Allen Lin, Lingfeng Shi, Ali Mostafavi, James Caverlee |
ACL (1) | 2 |
| 2026 | Multi-Scale Model Compression via Nested Matrix Learning
Xiangjue Dong, Aditya Anantharaman, Hemant Pugaliya |
LREC | 1 |
| 2026 | Language Models as Semantic Augmenters for Sequential RecommendersabstractLarge Language Models (LLMs) excel at capturing latent semantics and contextual relationships across diverse modalities. However, in modeling user behavior from sequential interaction data, performance often suffers when such semantic context is limited or absent. We introduce LaMAR, a LLM-driven semantic enrichment framework designed to enrich such sequences automatically. LaMAR leverages LLMs in a few-shot setting to generate auxiliary contextual signals by inferring latent semantic aspects of a user's intent and item relationships from existing metadata. These generated signals, such as inferred usage scenarios, item intents, or thematic summaries, augment the original sequences with greater contextual depth. We demonstrate the utility of this generated resource by integrating it into benchmark sequential modeling tasks, where it consistently improves performance. Further analysis shows that LLM-generated signals exhibit high semantic novelty and diversity, enhancing the representational capacity of the downstream models. This work represents a new data-centric paradigm where LLMs serve as intelligent context generators, contributing a new method for the semi-automatic creation of training data and language resources. Mahsa Valizadeh, Xiangjue Dong, Rui Tuo 0001, James Caverlee |
LREC | 2 |
| 2025 | Masculine Defaults via Gendered Discourse in Podcasts and Large Language ModelsabstractMasculine defaults are widely recognized as a significant type of gender bias, but they are often unseen as they are under-researched. Masculine defaults involve three key parts: (i) the cultural context, (ii) the masculine characteristics or behaviors, and (iii) the reward for, or simply acceptance of, those masculine characteristics or behaviors. In this work, we study discourse-based masculine defaults, and propose a twofold framework for (i) the large-scale discovery and analysis of gendered discourse words in spoken content via our Gendered Discourse Correlation Framework (GDCF); and (ii) the measurement of the gender bias associated with these gendered discourse words in LLMs via our Discourse Word-Embedding Association Test (D-WEAT). We focus our study on podcasts, a popular and growing form of social media, analyzing 15,117 podcast episodes. We analyze correlations between gender and discourse words -- discovered via LDA and BERTopic -- to automatically form gendered discourse word lists. We then study the prevalence of these gendered discourse words in domain-specific contexts, and find that gendered discourse-based masculine defaults exist in the domains of business, technology/politics, and video games. Next, we study the representation of these gendered discourse words from a state-of-the-art LLM embedding model from OpenAI, and find that the masculine discourse words have a more stable and robust representation than the feminine discourse words, which may result in better system performance on downstream tasks for men. Hence, men are rewarded for their discourse patterns with better system performance by one of the state-of-the-art language models -- and this embedding disparity is a representational harm and a masculine default. Maria Teleki, Xiangjue Dong, James Caverlee |
ICWSM | 2 |
| 2024 | DACL: Disfluency Augmented Curriculum Learning for Fluent Text GenerationabstractVoice-driven software systems are in abundance. However, language models that power these systems are traditionally trained on fluent, written text corpora. Hence there can be a misalignment between the inherent disfluency of transcribed spoken content and the fluency of the written training data. Furthermore, gold-standard disfluency annotations of various complexities for incremental training can be expensive to collect. So, we propose in this paper a Disfluency Augmented Curriculum Learning (DACL) approach to tackle the complex structure of disfluent sentences and generate fluent texts from them, by using Curriculum Learning (CL) coupled with our synthetically augmented disfluent texts of various levels. DACL harnesses the tiered structure of our generated synthetic disfluent data using CL, by training the model on basic samples (i.e. more fluent) first before training it on more complex samples (i.e. more disfluent). In contrast to the random data exposure paradigm, DACL focuses on a simple-to-complex learning process. We comprehensively evaluate DACL on Switchboard Penn Treebank-3 and compare it to the state-of-the-art disfluency removal models. Our model surpasses existing techniques in word-based precision (by up to 1%) and has shown favorable recall and F1 scores. Rohan Chaudhury, Maria Teleki, Xiangjue Dong, James Caverlee |
LREC/COLING | 3 |
| 2024 | Quantifying the Impact of Disfluency on Spoken Content SummarizationabstractSpoken content is abundant – including podcasts, meeting transcripts, and TikTok-like short videos. And yet, many important tasks like summarization are often designed for written content rather than the looser, noiser, and more disfluent style of spoken content. Hence, we aim in this paper to quantify the impact of disfluency on spoken content summarization. Do disfluencies negatively impact the quality of summaries generated by existing approaches? And if so, to what degree? Coupled with these goals, we also investigate two methods towards improving summarization in the presence of such disfluencies. We find that summarization quality does degrade with an increase in these disfluencies and that a combination of multiple disfluency types leads to even greater degradation. Further, our experimental results show that naively removing disfluencies and augmenting with special tags can worsen the summarization when used for testing, but that removing disfluencies for fine-tuning yields the best results. We make the code available at https://github.com/mariateleki/Quantifying-Impact-Disfluency. Maria Teleki, Xiangjue Dong, James Caverlee |
LREC/COLING | 2 |
| 2024 | The Neglected Tails in Vision-Language ModelsabstractVision-language models (VLMs) excel in zero-shot recognition but their performance varies greatly across different visual concepts. For example, although CLIP achieves impressive accuracy on ImageNet (60-80%), its performance drops below 10% for more than ten concepts like night snake, presumably due to their limited presence in the pretraining data. However, measuring the frequency of concepts in VLMs' large-scale datasets is challenging. We address this by using large language models (LLMs) to count the number of pretraining texts that con-tain synonyms of these concepts. Our analysis confirms that popular datasets, such as LAION, exhibit a long-tailed concept distribution, yielding biased performance in VLMs. We also find that downstream applications of VLMs, including visual chatbots (e.g., GPT-4V) and text-to-image models (e.g., Stable Diffusion), often fail to recognize or generate images of rare concepts identified by our method. To mit-igate the imbalanced performance of zero-shot VLMs, we propose REtrieval-Augmented Learning (REAL). First, in-stead of prompting VLMs using the original class names, REAL uses their most frequent synonyms found in pretraining texts. This simple change already outperforms costly human-engineered and LLM-enriched prompts over nine benchmark datasets. Second, REAL trains a linear classifier on a small yet balanced set of pretraining data re-trieved using concept synonyms. REAL surpasses the previous zero-shot SOTA, using 400× less storage and 10,000× less training time! Shubham Parashar, Zhiqiu Lin, Tian Liu 0006, Xiangjue Dong, Deva Ramanan, James Caverlee, Shu Kong |
CVPR | 4 |
| 2024 | DA³: A Distribution-Aware Adversarial Attack against Language ModelsabstractLanguage models can be manipulated by adversarial attacks, which introduce subtle perturbations to input data.While recent attack methods can achieve a relatively high attack success rate (ASR), we've observed that the generated adversarial examples have a different data distribution compared with the original examples.Specifically, these adversarial examples exhibit reduced confidence levels and greater divergence from the training data distribution.Consequently, they are easy to detect using straightforward detection methods, diminishing the efficacy of such attacks.To address this issue, we propose a Distribution-Aware Adversarial Attack (DA 3 ) method.DA 3 considers the distribution shifts of adversarial examples to improve attacks' effectiveness under detection methods.We further design a novel evaluation metric, the Non-detectable Attack Success Rate (NASR), which integrates both ASR and detectability for the attack task.We conduct experiments on four widely used datasets to validate the attack effectiveness and transferability of adversarial examples generated by DA 3 against both the white-box BERT-BASE and ROBERTA-BASE models and the black-box LLAMA2-7B model 1 . Yibo Wang 0001, Xiangjue Dong, James Caverlee, Philip S. Yu |
EMNLP | 2 |
| 2024 | Comparing ASR Systems in the Context of Speech Disfluencies
Maria Teleki, Xiangjue Dong, Soohwan Kim, James Caverlee |
INTERSPEECH | 2 |
| 2023 | Closed-book Question Generation via Contrastive LearningabstractQuestion Generation (QG) is a fundamental NLP task for many downstream applications.Recent studies on open-book QG, where supportive answer-context pairs are provided to models, have achieved promising progress.However, generating natural questions under a more practical closed-book setting that lacks these supporting documents still remains a challenge.In this work, we propose a new QG model for this closed-book setting that is designed to better understand the semantics of long-form abstractive answers and store more information in its parameters through contrastive learning and an answer reconstruction module.Through experiments, we validate the proposed QG model on both public datasets and a new WikiCQA dataset.Empirical results show that the proposed QG model outperforms baselines in both automatic evaluation and human evaluation.In addition, we show how to leverage the proposed model to improve existing question-answering systems.These results further indicate the effectiveness of our QG model for enhancing closed-book questionanswering tasks. Xiangjue Dong, Jiaying Lu 0001, Jianling Wang, James Caverlee |
EACL | 1 |
| 2023 | Weakly Supervised Concept Map Generation Through Task-Guided Graph TranslationabstractRecent years have witnessed the rapid development of concept map generation techniques due to their advantages in providing well-structured summarization of knowledge from free texts. Traditional unsupervised methods do not generate task-oriented concept maps, whereas deep generative models require large amounts of training data. In this work, we presentGT-D2G(Graph Translation-based Document To Graph), an automatic concept map generation framework that leverages generalized NLP pipelines to derive semantic-rich initial graphs, and translates them into more concise structures under the weak supervision of downstream task labels. The concept maps generated byGT-D2Gcan provide interpretable summarization of structured knowledge for the input texts, which are demonstrated through human evaluation and case studies on three real-world corpora. Further experiments on the downstream task of document classification show thatGT-D2Gbeats other concept map generation methods. Moreover, we specifically validate the labeling efficiency ofGT-D2Gin the label-efficient learning setting and the flexibility of generated graph sizes in controlled hyper-parameter studies. Jiaying Lu 0001, Xiangjue Dong, Carl Yang 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |