VLDB 2026 Research / reviewers in the wild / expert
Nan Xu 0004
dblp:26/4304-4
· DBLP profile ↗
23ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0002-9794-7262ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 4 first-author · 9 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 1 since 2021Security and privacy · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language AgentsabstractMinzheng Wang, Run Luo, Yanbo Wang, Zichen Liu, Yuqiao Tan, Tao Tan, Nan Xu, Lu Wang, Wenji Mao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Minzheng Wang 0001, Run Luo, Yanbo Wang 0004, Yuqiao Tan, Nan Xu 0004, Wenji Mao |
ACL (1) | 7 |
| 2026 | mChartQA and mChartQABench: A multimodal-only solution for complex chart question-answering
Jingxuan Wei, Nan Xu 0004, Guiyong Chang, Yin Luo, Bihui Yu, Ruifeng Guo |
Pattern Recognit. | 2 |
| 2024 | Bridging Word-Pair and Token-Level Metaphor Detection with Explainable Domain MiningabstractMetaphor detection aims to identify whether a linguistic expression in text is metaphorical or literal.Most existing research tackles this problem either using word-pair or token-level information as input, and thus treats word-pair and token-level metaphor detection as distinct subtasks.Benefited from the simplified structure of word pairs, recent methods for wordpair metaphor detection can provide intermediate explainable clues for the detection results, which remains a challenging issue for tokenlevel metaphor detection.To mitigate this issue in token-level metaphor detection and take advantage of word pairs, in this paper, we make the first attempt to bridge word-pair and tokenlevel metaphor detection via modeling word pairs within a sentence as explainable intermediate information.As the central role of verb in metaphorical expressions, we focus on tokenlevel verb metaphor detection and propose a novel explainable Word Pair based Domain Mining (WPDM) method.Our work is inspired by conceptual metaphor theory (CMT).We first devise an approach for conceptual domain mining utilizing semantic role mapping and resources at cognitive, commonsense and lexical levels.We then leverage the inconsistency between source and target domains for core word pair modeling to facilitate the explainability.Experiments on four datasets verify the effectiveness of our method and demonstrate its capability to provide the core word pair and corresponding conceptual domains as explainable clues for metaphor detection. Ruike Zhang, Nan Xu 0004, Wenji Mao |
ACL (1) | 3 |
| 2024 | PromISe: Releasing the Capabilities of LLMs with Prompt Introspective SearchabstractThe development of large language models (LLMs) raises the importance of assessing the fairness and completeness of various evaluation benchmarks. Regrettably, these benchmarks predominantly utilize uniform manual prompts, which may not fully capture the expansive capabilities of LLMs—potentially leading to an underestimation of their performance. To unlock the potential of LLMs, researchers pay attention to automated prompt search methods, which employ LLMs as optimizers to discover optimal prompts. However, previous methods generate the solutions implicitly, which overlook the underlying thought process and lack explicit feedback. In this paper, we propose a novel prompt introspective search framework, namely PromISe, to better release the capabilities of LLMs. It converts the process of optimizing prompts into an explicit chain of thought, through a step-by-step procedure that integrates self-introspect and self-refine. Extensive experiments, conducted over 73 tasks on two major benchmarks, demonstrate that our proposed PromISe significantly boosts the performance of 12 well-known LLMs compared to the baseline approach. Moreover, our study offers enhanced insights into the interaction between humans and LLMs, potentially serving as a foundation for future designs and implementations. Keywords: large language models, prompt search, self-introspect, self-refine Minzheng Wang 0001, Nan Xu 0004, Jiahao Zhao 0001, Yin Luo, Wenji Mao |
LREC/COLING | 2 |
| 2024 | MMA-Diffusion: MultiModal Attack on Diffusion ModelsabstractIn recent years, Text-to-Image (T2I) models have seen remarkable advancements, gaining widespread adoption. However, this progress has inadvertently opened avenues for potential misuse, particularly in generating inappropriate or Not-Safe-For-Work (NSFW) content. Our work introduces MMA-Diffusion, a framework that presents a significant and realistic threat to the security of T2I models by effectively circumventing current defensive measures in both open-source models and commercial online services. Unlike previous approaches, MMA-Diffusion leverages both textual and visual modalities to bypass safeguards like prompt filters and post-hoc safety checkers, thus exposing and highlighting the vulnerabilities in existing defense mechanisms. Our codes are available at https://github.com/cure-lab/MMA-Diffusion. Ruiyuan Gao 0001, Xiaosen Wang, Tsung-Yi Ho, Nan Xu 0004, Qiang Xu 0001 |
CVPR | 5 |
| 2024 | Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QAabstractMinzheng Wang, Longze Chen, Fu Cheng, Shengyi Liao, Xinghua Zhang, Bingli Wu, Haiyang Yu, Nan Xu, Lei Zhang, Run Luo, Yunshui Li, Min Yang, Fei Huang, Yongbin Li. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Minzheng Wang 0001, Longze Chen, Fu Cheng, Shengyi Liao, Xinghua Zhang 0001, Bingli Wu, Haiyang Yu 0003, Nan Xu 0004, Lei Zhang 0201, Run Luo, Yunshui Li, Min Yang 0007, Fei Huang 0002, Yongbin Li 0001 |
EMNLP | 8 |
| 2024 | A Theory Guided Scaffolding Instruction Framework for LLM-Enabled Metaphor ReasoningabstractYuan Tian, Nan Xu, Wenji Mao. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Nan Xu 0004, Wenji Mao |
NAACL-HLT | 2 |
| 2023 | Dynamic Routing Transformer Network for Multimodal Sarcasm DetectionabstractMultimodal sarcasm detection is an important research topic in natural language processing and multimedia computing, and benefits a wide range of applications in multiple domains.Most existing studies regard the incongruity between image and text as the indicative clue in identifying multimodal sarcasm.To capture cross-modal incongruity, previous methods rely on fixed architectures in network design, which restricts the model from dynamically adjusting to diverse imagetext pairs.Inspired by routing-based dynamic network, we model the dynamic mechanism in multimodal sarcasm detection and propose the Dynamic Routing Transformer Network (DynRT-Net).Our method utilizes dynamic paths to activate different routing transformer modules with hierarchical co-attention adapting to cross-modal incongruity.Experimental results on a public dataset demonstrate the effectiveness of our method compared to the stateof-the-art methods.Our codes are available at https://github.com/TIAN-viola/DynRT. Nan Xu 0004, Ruike Zhang, Wenji Mao |
ACL (1) | 2 |
| 2023 | Modeling Conceptual Attribute Likeness and Domain Inconsistency for Metaphor DetectionabstractMetaphor detection is an important and challenging task in natural language processing, which aims to distinguish between metaphorical and literal expressions in text.Previous studies mainly leverage the incongruity of source and target domains and contextual clues for detection, neglecting similar attributes shared between source and target concepts in metaphorical expressions.Based on conceptual metaphor theory, these similar attributes are essential to infer implicit meanings conveyed by the metaphor.Under the guidance of conceptual metaphor theory, in this paper, we model the likeness of attribute for the first time and propose a novel Attribute lIkeness and Domain Inconsistency Learning framework (AIDIL) for word-pair metaphor detection.Specifically, we propose an attribute siamese network to mine similar attributes between source and target concepts.We then devise a domain contrastive learning strategy to learn the semantic inconsistency of concepts in source and target domains.Extensive experiments on four datasets verify that our method significantly outperforms the previous state-of-the-art methods, and demonstrate the generalization ability of our method. Nan Xu 0004, Wenji Mao, Daniel Dajun Zeng |
EMNLP | 2 |
| 2023 | Adversarial Topic-Aware Memory Network for Cross-Lingual Stance DetectionabstractStance detection is an important research area in social media analytics and text mining and has a number of security-related applications, which aims to identify the user's viewpoint towards specific targets in domains such as social security, politics and public events. Previous research on stance detection has mainly focused on monolingual scenarios with the limited number of targets, while little attention was paid to cross-lingual stance detection. In contrast to the abundant labeled data in source language (typically English), the labeled data are often scarce in the non-English target language. Meanwhile, the diverse expressions of target further complicate the task and bring additional challenges. In this paper, we focus on cross-lingual stance detection in practical applications, where target is generally expressed as a topic and no labeled data are available for the target language. To tackle the above challenges, we propose an adversarial topic-aware memory network (ATOM) for cross-lingual stance detection. Specifically, our method first mines the generalized topic representations across source and target languages and utilizes them as the guidance to transfer knowledge from the high-resource source language to the low-resource target language. We further develop an iterative memory network to facilitate knowledge transfer across languages, which adaptively generates language-invariant topic-aware clues via adversarial training. Experimental results on three multilingual datasets in the politics domain demonstrate the effectiveness of our proposed method. Ruike Zhang, Nan Xu 0004, Wenji Mao, Daniel Dajun Zeng |
ISI | 2 |
| 2023 | Controllable News Comment Generation based on Attribute Level Contrastive LearningabstractNews comments provide a convenient way for people to express opinions and exchange ideas. Positive comments en-courage a harmonious discussion atmosphere within news media communities. In contrast, offensive or insulting comments may result in cyberbullying and personal psychological trauma, which have particular practical impacts in security-related domains. The automatic generation of news comments with controllable attributes (e.g. sentiment) to assist users and news platform administrators is greatly needed. However, existing research for news comment generation has not addressed the controllable issue yet. On the other hand, existing methods for controllable text generation focus on token-level constraints, which are not applicable to controlling the sentence level attributes for news comment generation. To address this challenging issue, in this paper, we propose an attribute-level contrastive learning method for controllable news comment generation. To apply attribute level constraints on the generated text, our method considers the attributes of the generated comments and the pre-defined attributes as different views of the same attribute, and maximizes their similarity during the training process. We conduct experiments on two publicly available news comment datasets, and the experimental results show that our model achieves competitive performance in news comment generation and attribute controllability. Hanyi Zou, Nan Xu 0004, Qingchao Kong, Wenji Mao |
ISI | 2 |
| 2023 | AnANet: Association and Alignment Network for Modeling Implicit Relevance in Cross-Modal Correlation ClassificationabstractWith the explosive increase of multimodal data, cross-modal correlation classification has become an important research topic and is in great demand in many cross-modal applications. A variety of classification schemes and predictive models have been built based on the existing cross-modal correlation categorization. However, these classification schemes typically follow the prior assumption that the paired cross-modal samples are strictly related, and thus pay great attention to the fine-grained relevant types of cross-modal correlation, ignoring the high volume of implicitly relevant data which are often wrongly classified into irrelevant types. Even more, previous predictive models fall short of reflecting the essence of cross-modal correlation according to their definitions, especially in the modeling of network structure. Thus in this paper, by comprehensively investigating the current image-text correlation classification research, we redefine a new classification scheme for cross-modal correlation based on the implicit and explicit relevance. To predict the types of image-text correlation based on our proposed definition, we further devise the Association and Alignment Network (namely AnANet) to model the implicit and explicit relevance, which captures both the implicit association of global discrepancy and commonality between image and text and explicit alignment of cross-modal local relevance. Experimental studies on our constructed new image-text correlation dataset verify the effectiveness of our proposed model. Nan Xu 0004, Ruike Zhang, Wenji Mao |
IEEE Trans. Multim. | 1 |
| 2021 | A Central Opinion Extraction Framework for Boosting Performance on Sentiment AnalysisabstractWith the rapid development of the Internet, mining opinions and emotions from the explosive growth of user-generated content is a key field of social media analysis. However, the expression forms of the central opinion which strongly expresses the essential points and converges the main sentiments of the overall document are diverse in practice, such as sequential sentences, a sentence fragment, or an individual sentence. Previous research studies on sentiment analysis based on document level and sentence level fail to deal with this actual situation uniformly. To address this issue, we propose a Central Opinion Extraction (COE) framework to boost performance on sentiment analysis with social media texts. Our framework first extracts a span-level central opinion text, which expresses the essential opinion related to sentiment representation among the whole text, and then uses extracted textual span to boost the performance of sentiment classifiers. The experimental results on a public dataset show the effectiveness of our framework for boosting the performance on document-level sentiment analysis task. Nan Xu 0004, Wenji Mao, Yin Luo |
ISI | 2 |
| 2021 | PAN: Prototype-based Adaptive Network for Robust Cross-modal RetrievalabstractIn practical applications of cross-modal retrieval, test queries of the retrieval system may vary greatly and come from unknown category. Meanwhile, due to the cost and difficulty of data collection as well as other issues, the available data for cross-modal retrieval are often imbalanced over different modalities. In this paper, we address two important issues to increase the robustness of cross-modal retrieval system for real-world applications: handling test queries from unknown category and modality-imbalanced training data. The first issue has not been addressed by existing methods and the second issue was not well addressed in the related research. To tackle the above issues, we take the advantage of prototype learning, and propose a prototype-based adaptive network (PAN) for robust cross-modal retrieval. Our method leverages a unified prototype to represent each semantic category across modalities, which provides discriminative information of different categories and takes unified prototypes as anchors to learn cross-modal representations adaptively. Moreover, we propose a novel prototype propagation strategy to reconstruct balanced representations which preserves the semantic consistency and modality heterogeneity. Experimental results on the benchmark datasets demonstrate the effectiveness of our method compared to the SOTA methods, and further robustness tests show the superiority of our method in solving the above issues. Zhixiong Zeng, Nan Xu 0004, Wenji Mao |
SIGIR | 3 |
| 2020 | Reasoning with Multimodal Sarcastic Tweets via Modeling Cross-Modality Contrast and Semantic AssociationabstractSarcasm is a sophisticated linguistic phenomenon to express the opposite of what one really means.With the rapid growth of social media, multimodal sarcastic tweets are widely posted on various social platforms.In multimodal context, sarcasm is no longer a pure linguistic phenomenon, and due to the nature of social media short text, the opposite is more often manifested via cross-modality expressions.Thus traditional text-based methods are insufficient to detect multimodal sarcasm.To reason with multimodal sarcastic tweets, in this paper, we propose a novel method for modeling cross-modality contrast in the associated context.Our method models both cross-modality contrast and semantic association by constructing the Decomposition and Relation Network (namely D&R Net).The decomposition network represents the commonality and discrepancy between image and text, and the relation network models the semantic association in cross-modality context.Experimental results on a public dataset demonstrate the effectiveness of our model in multimodal sarcasm detection. Nan Xu 0004, Zhixiong Zeng, Wenji Mao |
ACL | 1 |
| 2020 | Event-Driven Network for Cross-Modal RetrievalabstractDespite extensive research on cross-modal retrieval, existing methods focus on the matching between image objects and text words. However, for the large amount of social media, such as news reports and online posts with images, previous methods are insufficient to model the associations between long text and image. As long text contains multiple entities and relationships between them, as well as complex events sharing a common scenario of the text, it poses unique research challenge to cross-modal retrieval. To tackle the challenge, in this paper, we focus on the retrieval task on long text and image, and propose an event-driven network for cross-modal retrieval. Our approach consists of two modules, namely the contextual neural tensor network (CNTN) and cross-modal matching network (CMMN). The CNTN module captures both event-level and text-level semantics of the sequential events extracted from a long text. The CMMN module learns a common representation space to compute the similarity of image and text modalities. We construct a multimodal dataset based on the news reports in People's Daily. The experimental results demonstrate that our model outperforms the existing state-of-the-art methods and can provide semantic richer text representations to enhance the effectiveness in cross-modal retrieval. Zhixiong Zeng, Nan Xu 0004, Wenji Mao |
CIKM | 2 |
| 2019 | Multi-Interactive Memory Network for Aspect Based Multimodal Sentiment AnalysisabstractAs a fundamental task of sentiment analysis, aspect-level sentiment analysis aims to identify the sentiment polarity of a specific aspect in the context. Previous work on aspect-level sentiment analysis is text-based. With the prevalence of multimodal user-generated content (e.g. text and image) on the Internet, multimodal sentiment analysis has attracted increasing research attention in recent years. In the context of aspect-level sentiment analysis, multimodal data are often more important than text-only data, and have various correlations including impacts that aspect brings to text and image as well as the interactions associated with text and image. However, there has not been any related work carried out so far at the intersection of aspect-level and multimodal sentiment analysis. To fill this gap, we are among the first to put forward the new task, aspect based multimodal sentiment analysis, and propose a novel Multi-Interactive Memory Network (MIMN) model for this task. Our model includes two interactive memory networks to supervise the textual and visual information with the given aspect, and learns not only the interactive influences between cross-modality data but also the self influences in single-modality data. We provide a new publicly available multimodal aspect-level sentiment dataset to evaluate our model, and the experimental results demonstrate the effectiveness of our proposed model for this new task. Nan Xu 0004, Wenji Mao, Guandan Chen |
AAAI | 1 |
| 2019 | Modeling Conversation Structure and Temporal Dynamics for Jointly Predicting Rumor Stance and VeracityabstractPenghui Wei, Nan Xu, Wenji Mao. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Penghui Wei, Nan Xu 0004, Wenji Mao |
EMNLP/IJCNLP (1) | 2 |
| 2019 | NPP: A neural popularity prediction model for social media content
Guandan Chen, Qingchao Kong, Nan Xu 0004, Wenji Mao |
Neurocomputing | 3 |
| 2018 | An Encoder-Memory-Decoder Framework for Sub-Event Detection in Social MediaabstractSub-event detection can help faster and deeper understanding of an event by providing human-friendly clusters, and thus has become an important research topic in Web mining and knowledge management. In existing sub-event detection methods, clustering based methods are brittle for using heuristic similarity metric to judge whether documents belong to the same sub-event, while topic model based methods are limited to the bag of words assumption. To overcome these drawbacks in previous research, in this paper, we propose an encoder-memory-decoder framework for sub-event detection. Our model learns document and sub-event representations suitable for the similarity metric in a data-driven manner, and transforms sub-event detection into selecting the most proper sub-event representation that can maximize text reconstruction probability. Considering the case of over-fitting, we also apply transfer learning in our model. To the best of our knowledge, our model is the first to develop an unsupervised deep neural model for sub-event detection. We use Twitter as an examplar social media platform for our study, and experimental results show that our model outperforms baseline methods for sub-event detection. Guandan Chen, Nan Xu 0004, Wenji Mao |
CIKM | 2 |
| 2018 | MNRD: A Merged Neural Model For Rumor Detection In Social MediaabstractWith the continuous development of social media and the diversity of its contents, rumors can quickly spread widely and cause serious damage, which makes the early detection of rumors a key task in social media analysis. Most existing methods for rumor detection exploit handcrafted features and ignore the different roles original post and retweets play in information production and spreading process. In fact, original post usually contains detailed content information about event itself, while retweets reflect the diffusion of information in online interaction. Recently, several studies manifest the effectiveness of deep neural network in rumor detection. However, these previous research has made no distinction between original post and retweets for rumor detection. In this paper, we propose a Merged Neural Rumor Detection (MNRD) model that consider three aspects of rumor data: content of original post, diffusion of retweet and user information. We propose content attention to aggregate pivotal words in original post and diffusion attention to focus on most informative retweets during the diffusion process. We also incorporate user information by identifying features to capture user's reliability and social influence. Experimental results on real world dataset show the effectiveness of our model to detect rumors in social media. Nan Xu 0004, Guandan Chen, Wenji Mao |
IJCNN | 1 |
| 2018 | A Co-Memory Network for Multimodal Sentiment AnalysisabstractWith the rapid increase of diversity and modality of data in user-generated contents, sentiment analysis as a core area of social media analytics has gone beyond traditional text-based analysis. Multimodal sentiment analysis has become an important research topic in recent years. Most of the existing work on multimodal sentiment analysis extracts features from image and text separately, and directly combine them to train a classifier. As visual and textual information in multimodal data can mutually reinforce and complement each other in analyzing the sentiment of people, previous research all ignores this mutual influence between image and text. To fill this gap, in this paper, we consider the interrelation of visual and textual information, and propose a novel co-memory network to iteratively model the interactions between visual contents and textual words for multimodal sentiment analysis. Experimental results on two public multimodal sentiment datasets demonstrate the effectiveness of our proposed model compared to the state-of-the-art methods. Nan Xu 0004, Wenji Mao, Guandan Chen |
SIGIR | 1 |
| 2017 | MultiSentiNet: A Deep Semantic Network for Multimodal Sentiment AnalysisabstractWith the prevalence of more diverse and multiform user-generated content in social networking sites, multimodal sentiment analysis has become an increasingly important research topic in recent years. Previous work on multimodal sentiment analysis directly extracts feature representation of each modality and fuse these features for classification. Consequently, some detailed semantic information for sentiment analysis and the correlation between image and text have been ignored. In this paper, we propose a deep semantic network, namely MultiSentiNet, for multimodal sentiment analysis. We first identify object and scene as salient detectors to extract deep semantic features of images. We then propose a visual feature guided attention LSTM model to extract words that are important to understand the sentiment of whole tweet and aggregate the representation of those informative words with visual semantic features, object and scene. The experiments on two public available sentiment datasets verify the effectiveness of our MultiSentiNet model and show that our extracted semantic features demonstrate high correlations with human sentiments. Nan Xu 0004, Wenji Mao |
CIKM | 1 |