VLDB 2026 Research / reviewers in the wild / expert
Cong Xu 0001
dblp:47/4804-1
· DBLP profile ↗
12ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0002-9288-1743ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Conditional Information Bottleneck for Multimodal Fusion: Overcoming Shortcut Learning in Sarcasm DetectionabstractMultimodal sarcasm detection is a complex task that requires distinguishing subtle complementary signals across modalities while filtering out irrelevant information. Many advanced methods rely on learning shortcuts from datasets rather than extracting intended sarcasm-related features. However, our experiments show that shortcut learning impairs the model's generalization in real-world scenarios. Furthermore, we reveal the weaknesses of current modality fusion strategies for multimodal sarcasm detection through systematic experiments, highlighting the necessity of focusing on effective modality fusion for complex emotion recognition. To address these challenges, we construct MUStARD++R by removing shortcut signals from MUStARD++. Then, a Multimodal Conditional Information Bottleneck (MCIB) model is introduced to enable efficient multimodal fusion for sarcasm detection. Experimental results show that the MCIB achieves the best performance without relying on shortcut learning. Qi Jia 0004, Cong Xu 0001, Feiyu Chen 0005, Yuhan Liu 0014, Haotian Zhang 0017, Lu Liu 0009, Zhichun Wang |
AAAI | 3 |
| 2026 | T-MSA: Transformer-Driven Multi-Strategy Adaptive Microarchitecture Design Space ExplorationabstractThe design of modern processors ignores the topological relationships among all design parameters, leading to significant simulation costs wasted on invalid designs. Therefore, we propose the T-MSA to address this issue. It is a Transformer-driven multi-strategy adaptive design space exploration scheme. A customized lightweight Transformer (LiteFormer) is devised to model topological relationships among arbitrary design parameters, constructing an implicit interaction graph in the latent space. Secondly, we design a dynamic active learning (DynamicAL) strategy to extract sparse and high-quality initial points via sparse centroid initialization and hybrid sampling. Finally, a triple Pareto frontier acquisition function (TriPFAF) is devised to guide optimization direction based on gains from three types of Pareto frontiers, dynamically balancing exploration and exploitation. We conducted rigorous experiments on two BOOM evaluation platforms, demonstrating that T-MSA efficiently and comprehensively optimizes the performance-power-area (PPA) objective. The designs it identifies achieve significant improvements over state-of-the-art DSE algorithms on Pareto hypervolume (HV). When attaining the same HV value, T-MSA outperforms BOOM-Explorer by 188.24% and 133.33% on two platforms. Fan Yang 0032, Xiaochuan Li 0001, Cong Xu 0001, RenGang Li, Baoyu Fan |
DATE | 6 |
| 2026 | Visual Question Explainable Reasoning on Hypothesis Agent Interaction with Scene
Baoyu Fan, Cong Xu 0001, Lu Liu 0009, Xiaoli Gong, Jin Zhang 0003 |
Signal Process. | 2 |
| 2025 | Dropletvideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video Generation
Guoguang Du 0001, Xiaochuan Li 0001, Qi Jia 0004, Lu Liu 0009, Cong Xu 0001, Zhenhua Guo 0003, Yaqian Zhao, Xiaoli Gong, RenGang Li, Baoyu Fan |
ICCV | 8 |
| 2025 | A Multi-Granularity Relation Graph Aggregation Framework With Multimodal Clues for Social Relation ReasoningabstractThe social relation is a fundamental attribute of human beings in daily life. The ability of humans to form large organizations and institutions stems directly from our complex social networks. Therefore, understanding social relationships in the context of multimedia is crucial for building domain-specific or general artificial intelligence systems. The key to reason social relations lies in understanding the human interactions between individuals through multimodal representations such as action and utterance. However, due to video editing techniques and various narrative sequences in videos, two individuals with social relationships may not appear together in the same frame or clip. Additionally, social relations may manifest in different levels of granularity in video expressions. Previous research has not effectively addressed these challenges. Therefore, this paper proposes aMulti-Granularity Relation Graph Aggregation Framework(MGRG) to enhance the inference ability for social relation reasoning in multimedia content, like video. Different from existing methods, our method considers the paradigm of jointly inferring the relations by constructing a social relation graph. We design a hierarchical multimodal relation graph illustrating the exchange of information between individuals' roles, capturing the complex interactions at multi-levels of granularity from fine to coarse. In MGRG, we propose two aggregation modules to cluster multimodal features in different granularity layer relation graph, considering temporal aspects and importance. Experimental results show that our method generates a logical and coherent social relation graph and improves the performance in accuracy. Cong Xu 0001, Feiyu Chen 0005, Qi Jia 0004, Yunji Li, Yaqian Zhao, Changming Zhao |
IEEE Trans. Multim. | 1 |
| 2024 | Egocentric Vehicle Dense Video CaptioningabstractTraditional dense video captioning predominantly focuses on edited exocentric footage. These videos are filmed from an external perspective and generally feature distinct transitions between different events, as exemplified in edited instructional videos. However, such videos do not genuinely reflect the way we perceive our real lives. Instead, we observe the world from an egocentric viewpoint and witness only continuous unedited footage. To facilitate further research, we introduce a new topic: Egocentric Vehicle Dense Video Captioning, in classic vehicle driving scenarios. This is a multi-modal, multi-task subject endeavor for a comprehensive understanding of untrimmed, egocentric driving videos. It consists of three sub-tasks that concentrate on event localization, captioning, and vehicle state estimation separately. To accomplish these tasks, it is necessary to deal with at least three challenges: extracting ego-motion relevant information, describing driving behavior and analyzing the underlying rationale, as well as resolving the boundary ambiguity problem. In response, we devise corresponding solutions, including a vehicle ego-motion learning strategy and a novel adjacent contrastive learning strategy, which effectively address the aforementioned issues. We validate our method by conducting extensive experiments on the BDD-X dataset, all of which show promising results and achieve new state-of-the-art performance on most metrics, which proves the effect of our approach. Feiyu Chen 0005, Cong Xu 0001, Qi Jia 0004, Yuhan Liu 0014, Haotian Zhang 0017, Endong Wang |
ACM Multimedia | 2 |
| 2024 | Infer Induced Sentiment of Comment Response to Video: A New Task, Dataset and BaselineabstractExisting video multi-modal sentiment analysis mainly focuses on the sentiment expression of people within the video, yet often neglects the induced sentiment of viewers while watching the videos. Induced sentiment of viewers is essential for inferring the public response to videos and has broad application in analyzing public societal sentiment, effectiveness of advertising and other areas. The micro videos and the related comments provide a rich application scenario for viewers’ induced sentiment analysis. In light of this, we introduces a novel research task, Multimodal Sentiment Analysis for Comment Response of Video Induced(MSA-CRVI), aims to infer opinions and emotions according to comments response to micro video. Meanwhile, we manually annotate a dataset named Comment Sentiment toward to Micro Video (CSMV) to support this research. It is the largest video multi-modal sentiment dataset in terms of scale and video duration to our knowledge, containing 107, 267 comments and 8, 210 micro videos with a video duration of 68.83 hours. To infer the induced sentiment of comment should leverage the video content, we propose the Video Content-aware Comment Sentiment Analysis (VC-CSA) method as a baseline to address the challenges inherent in this new task. Extensive experiments demonstrate that our method is showing significant improvements over other established baselines. We make the dataset and source code publicly available at https://github.com/IEIT-AGI/MSA-CRVI. Qi Jia 0004, Baoyu Fan, Cong Xu 0001, Lu Liu 0009, Guoguang Du 0001, Zhenhua Guo 0003, Yaqian Zhao, Xuanjing Huang 0001, RenGang Li |
NeurIPS | 3 |
| 2024 | MVIndEmo: a dataset for micro video public-induced emotion prediction on social mediaabstractAbstract Distinct from the realm of perceived emotion research, induced emotion pertains to the emotional responses engendered within content consumers. This facet has garnered considerable attention and finds extensive application in the analysis of public social media. However, the advent of micro videos presents unique challenges when attempting to discern the induced emotional patterns exhibited by content consumers, owing to their free-style representation and other factors. Consequently, we have put forth two novel tasks concerning the recognition of public-induced emotion on micro videos: emotion polarity and emotion classification. Additionally, we have introduced a accessible dataset specifically tailored for the analysis of public-induced emotion on micro videos. The data corpus has been meticulously collected from Tiktok, a burgeoning social media platform renowned for its trendsetting content. To construct the dataset, we have selected eight captivating topics that elicit vibrant social discussions. In devising our label generation strategy, we have employed an automated approach characterized by the fusion of multiple expert models. This strategy incorporates a confidence measure method that relies on three distinct models for effectively aggregating user comments. To accommodate adaptable benchmark configurations, we provide both binary classification labels and probability distribution labels. The dataset encompasses a vast collection of 7,153 labeled micro videos. We have undertaken an extensive statistical analysis of the dataset to provide a comprehensive overview composition. It is our earnest aspiration that this dataset will serve as a catalyst for pioneering research avenues in the analysis of emotional patterns and the understanding of multi-modal information. Zhenhua Guo 0003, Qi Jia 0004, Baoyu Fan, Cong Xu 0001, Yaqian Zhao, RenGang Li |
Multim. Syst. | 5 |
| 2022 | AI-VQA: Visual Question Answering based on Agent Interaction with InterpretabilityabstractVisual Question Answering (VQA) serves as a proxy for evaluating the scene understanding of an intelligent agent by answering questions about images. Most VQA benchmarks to date are focused on those questions that can be answered through understanding visual content in the scene, such as simple counting, visual attributes, and even a little challenging questions that require extra encyclopedic knowledge. However, humans have a remarkable capacity to reason dynamic interaction on the scene, which is beyond the literal content of an image and has not been investigated so far. In this paper, we propose Agent Interaction Visual Question Answering (AI-VQA), a task investigating deep scene understanding if the agent takes a certain action. For this task, a model not only needs to answer action-related questions but also to locate the objects in which the interaction occurs for guaranteeing it truly comprehends the action. Accordingly, we make a new dataset based on Visual Genome and ATOMIC knowledge graph, including more than 19,000 manually annotated questions, and will make it publicly available. Besides, we also provide an annotation of the reasoning path while developing the answer for each question. Based on the dataset, we further propose a novel method, called ARE, that can comprehend the interaction and explain the reason based on a given event knowledge base. Experimental results show that our proposed method outperforms the baseline by a clear margin. RenGang Li, Cong Xu 0001, Zhenhua Guo 0003, Baoyu Fan, Yaqian Zhao, Weifeng Gong, Endong Wang |
ACM Multimedia | 2 |
| 2021 | Traditional Chinese medicine symptom normalization approach leveraging hierarchical semantic information and text matching with attention mechanism
Qi Jia 0004, Dezheng Zhang 0001, Shibing Yang, Yingjie Shi, Hu Tao, Cong Xu 0001, Xiong Luo, Yuekun Ma, Yonghong Xie |
J. Biomed. Informatics | 7 |
| 2020 | A model with length-variable attention for spoken language understanding
Cong Xu 0001, Qing Li 0015, Dezheng Zhang 0001, Jiarui Cui 0001, Zhenqi Sun |
Neurocomputing | 1 |
| 2020 | Deep successor feature learning for text generation
Cong Xu 0001, Qing Li 0015, Dezheng Zhang 0001, Yonghong Xie, Xisheng Li |
Neurocomputing | 1 |