EDBT 2026 Demo / reviewers in the wild / expert
Xiyuan Gao
dblp:289/4064
· DBLP profile ↗
14ranked-venue papers
6as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reconcile Gradient Modulation for Harmony Multimodal LearningabstractMultimodal learning frequently faces two coupled challenges: modality imbalance, where dominant modalities suppress others during training, and modality conflict, where opposing gradient directions hinder optimization. Existing methods typically address these issues in isolation, yet they are intrinsically correlated and most fundamentally reflected in the gradient space—severe imbalance may obscure conflicts, while suppressing conflict may homogenize features and worsen imbalance, affecting fusion performance. To jointly address this coupled challenge, we propose Reconcile Gradient Modulation (RGM), a unified framework that adaptively adjusts gradient magnitude and direction for harmony multimodal learning. The core of RGM is SynOrth Grad, which minimizes Dirichlet energy to perform minimal-gradient surgery. It enhances cooperation synergy when modalities are aligned and enforces orthogonality to preserve uniqueness in conflict situations, thus promoting stable and balanced learning. To guide this modulation, we propose Cumulative Gradient Energy (CGE) as a convergence-guaranteed measure of modality-wise progress, and construct a Balance-nonConflict Plane (BCP) for real-time diagnosis and control of training dynamics. Experiments on diverse benchmarks validate our effectiveness and generalizability, consistently outperforming counterparts that are designed to handle multimodal imbalance or conflict independently. Xiyuan Gao, Bing Cao 0002, Baoquan Gong, Pengfei Zhu 0001 |
AAAI | 1 |
| 2026 | AMuSeD: An Attentive Deep Neural Network for Multimodal Sarcasm Detection Incorporating Bimodal Data AugmentationabstractDetecting sarcasm effectively requires a nuanced understanding of context, including vocal tones and facial expressions. The progression towards multimodal computational methods in sarcasm detection, however, faces challenges due to the scarcity of data. To address this, we present AMuSeD (Attentive deep neural network for MUltimodal Sarcasm dEtection incorporating bi-modal Data augmentation). This approach utilizes the Multimodal Sarcasm Detection Dataset (MUStARD) and introduces a two-phase bimodal data augmentation strategy. The first phase involves generating varied text samples through Back-Translation from several secondary languages. The second phase involves the refinement of a FastSpeech2-based speech synthesis system, tailored specifically for sarcasm to retain sarcastic intonations. Alongside a cloud-based Text-to-Speech (TTS) service, this Fine-tuned FastSpeech2 system produces corresponding audio for the text augmentations. We also evaluate various attention mechanisms for selectively enhancing sarcasm-relevant features, finding self-attention to be the most efficient. Our experiments reveal that the proposed approach achieves a significant F1-score of 81.0% in text-audio modalities, surpassing even models that use three modalities from the MUStARD dataset. Xiyuan Gao, Shubhi Bansal, Kushaan Gowda, Shekhar Nayak, Nagendra Kumar 0001, Matt Coler |
IEEE Trans. Affect. Comput. | 1 |
| 2025 | Asymmetric Reinforcing Against Multi-Modal Representation BiasabstractThe strength of multimodal learning lies in its ability to integrate information from various sources, providing rich and comprehensive insights. However, in real-world scenarios, multi-modal systems often face the challenge of dynamic modality contributions, the dominance of different modalities may change with the environments, leading to suboptimal performance in multimodal learning. Current methods mainly enhance weak modalities to balance multimodal representation bias, which inevitably optimizes from a partialmodality perspective, easily leading to performance descending for dominant modalities. To address this problem, we propose an Asymmetric Reinforcing method against Multimodal representation bias (ARM). Our ARM dynamically reinforces the weak modalities while maintaining the ability to represent dominant modalities through conditional mutual information. Moreover, we provide an in-depth analysis that optimizing certain modalities could cause information loss and prevent leveraging the full advantages of multimodal data. By exploring the dominance and narrowing the contribution gaps between modalities, we have significantly improved the performance of multimodal learning, making notable progress in mitigating imbalanced multimodal learning. Xiyuan Gao, Bing Cao 0002, Pengfei Zhu 0001, Nannan Wang 0001, Qinghua Hu |
AAAI | 1 |
| 2025 | Intra-modal Relation and Emotional Incongruity Learning using Graph Attention Networks for Multimodal Sarcasm DetectionabstractSarcasm detection poses unique challenges due to the complex nature of sarcastic expressions often embedded across multiple modalities. Current methods frequently fall short in capturing the incongruent emotional cues that are essential for identifying sarcasm in multimodal contexts. In this paper, we present a novel method to capture the pair-wise emotional incongruities between modalities through a cross-modal Contrastive Attention Mechanism (CAM), leveraging advanced data augmentation techniques to enhance data diversity and Supervised Contrastive Learning (SCL) to obtain discriminative embeddings. Additionally, we employ Graph Attention Networks (GATs) to construct modality-specific graphs, capturing intra-modal dependencies. Experiments conducted on the MUStARD++ dataset demonstrate the efficacy of our approach, achieving a macro F1 score of 74.96%, which outperforms state-of-the-art methods. Devraj Raghuvanshi, Xiyuan Gao, Shubhi Bansal, Matt Coler, Nagendra Kumar 0001, Shekhar Nayak |
ICASSP | 2 |
| 2025 | A Multimodal Chinese Dataset for Cross-lingual Sarcasm DetectionabstractSarcasm is expressed through subtle cues like pitch, speech rate, and facial expressions, with patterns varying across languages, e.g., English speakers lower the pitch while Cantonese speakers raise it. While humans readily interpret these signals, computational models struggle, creating challenges for Human-Machine Interaction. Most multimodal sarcasm recognition research focuses on English and the lack of high-quality datasets for other languages hinders cross-lingual and cross-cultural studies. We introduce the Multimodal Chinese Sarcasm Dataset (MCSD), containing 10.57 hours of video. We propose a standardized annotation framework that captures annotator certainty to reflect the subjectivity of sarcasm, achieving a Fleiss'kappa of 0.74 (unweighted) and 0.79 (certainty-weighted). Validation of our dataset using SVM achieves a 76.64% F1-score in sarcasm detection. MCSD lays the foundation for robust cross-lingual sarcasm detection, contributing to advanced, human-centric systems. Xiyuan Gao, Bruce Xiao Wang, Shuming Huang, Shekhar Nayak, Matt Coler |
INTERSPEECH | 1 |
| 2025 | Leveraging Large Language Models for Sarcastic Speech Annotation in Sarcasm DetectionabstractSarcasm fundamentally alters meaning through tone and context, yet detecting it in speech remains a challenge due to data scarcity. In addition, existing detection systems often rely on multimodal data, limiting their applicability in contexts where only speech is available. To address this, we propose an annotation pipeline that leverages large language models (LLMs) to generate a sarcasm dataset. Using a publicly available sarcasm-focused podcast, we employ GPT-4o and LLaMA 3 for initial sarcasm annotations, followed by human verification to resolve disagreements. We validate this approach by comparing annotation quality and detection performance on a publicly available sarcasm dataset using a collaborative gating architecture. Finally, we introduce PodSarc, a large-scale sarcastic speech dataset created through this pipeline. The detection model achieves a 73.63% F1 score, demonstrating the dataset's potential as a benchmark for sarcasm detection research. Yuqing Zhang 0003, Xiyuan Gao, Shekhar Nayak, Matt Coler |
INTERSPEECH | 3 |
| 2025 | Multimodal Negative LearningabstractMultimodal learning systems often encounter challenges related to modality imbalance, where a dominant modality may overshadow others, thereby hindering the learning of weak modalities. Conventional approaches often force weak modalities to align with dominant ones in "Learning to be (the same)" (Positive Learning), which risks suppressing the unique information inherent in the weak modalities. To address this challenge, we offer a new learning paradigm: "Learning Not to be" (Negative Learning). Instead of enhancing weak modalities’ target-class predictions, the dominant modalities dynamically guide the weak modality to suppress non-target classes. This stabilizes the decision space and preserves modality-specific information, allowing weak modalities to preserve unique information without being over-aligned. We proceed to reveal the multimodal learning from a robustness perspective and theoretically derive the Multimodal Negative Learning (MNL) framework, which introduces a dynamic guidance mechanism tailored for negative learning. Our method provably tightens the robustness lower bound of multimodal learning by increasing the Unimodal Confidence Margin (UCoM) and reduces the empirical error of weak modalities, particularly under noisy and imbalanced scenarios. Extensive experiments across multiple benchmarks demonstrate the effectiveness and generalizability of our approach against the competing methods. The code will be available at: https://github.com/BaoquanGong/Multimodal-Negative-Learning.git Baoquan Gong, Xiyuan Gao, Pengfei Zhu 0001, Qinghua Hu, Bing Cao 0002 |
NeurIPS | 2 |
| 2025 | Spoken in Jest, Detected in Earnest: A Systematic Review of Sarcasm Recognition - Multimodal Fusion, Challenges, and Future ProspectsabstractSarcasm, a common feature of human communication, poses challenges in interpersonal interactions and human-machine interactions. Linguistic research has highlighted the importance of prosodic cues, such as variations in pitch, speaking rate, and intonation, in conveying sarcastic intent. Although previous work has focused on text-based sarcasm detection, the role of speech data in recognizing sarcasm has been underexplored. Recent advancements in speech technology emphasize the growing importance of leveraging speech data for automatic sarcasm recognition, which can enhance social interactions for individuals with neurodegenerative conditions and improve machine understanding of complex human language use, leading to more nuanced interactions. This systematic review is the first to focus on speech-based sarcasm recognition, charting the evolution from unimodal to multimodal approaches. It covers datasets, feature extraction, and classification methods, and aims to bridge gaps across diverse research domains. The findings include limitations in datasets for sarcasm recognition in speech, the evolution of feature extraction techniques from traditional acoustic features to deep learning-based representations, and the progression of classification methods from unimodal approaches to multimodal fusion techniques. In so doing, we identify the need for greater emphasis on cross-cultural and multilingual sarcasm recognition, as well as the importance of addressing sarcasm as a multimodal phenomenon, rather than a text-based challenge. Xiyuan Gao, Shekhar Nayak, Matt Coler |
IEEE Trans. Affect. Comput. | 1 |
| 2024 | A Functional Trade-off between Prosodic and Semantic Cues in Conveying SarcasmabstractThis study investigates the acoustic features of sarcasm and disentangles the interplay between the propensity of an utterance being used sarcastically and the presence of prosodic cues signaling sarcasm. Using a dataset of sarcastic utterances compiled from television shows, we analyze the prosodic features within utterances and key phrases belonging to three distinct sarcasm categories (embedded, propositional, and illocutionary), which vary in the degree of semantic cues present, and compare them to neutral expressions. Results show that in phrases where the sarcastic meaning is salient from the semantics, the prosodic cues are less relevant than when the sarcastic meaning is not evident from the semantics, suggesting a trade-off between prosodic and semantic cues of sarcasm at the phrase level. These findings highlight a lessened reliance on prosodic modulation in semantically dense sarcastic expressions and a nuanced interaction that shapes the communication of sarcastic intent. Xiyuan Gao, Yuqing Zhang 0003, Shekhar Nayak, Matt Coler |
INTERSPEECH | 2 |
| 2024 | Contextual Learning in Fourier Complex Field for VHR Remote Sensing ImagesabstractVery high-resolution (VHR) remote sensing (RS) image classification is the fundamental task for RS image analysis and understanding. Recently, Transformer-based models demonstrated outstanding potential for learning high-order contextual relationships from natural images with general resolution ( pixels) and achieved remarkable results on general image classification tasks. However, the complexity of the naive Transformer grows quadratically with the increase in image size, which prevents Transformer-based models from VHR RS image ( pixels) classification and other computationally expensive downstream tasks. To this end, we propose to decompose the expensive self-attention (SA) into real and imaginary parts via discrete Fourier transform (DFT) and, therefore, propose an efficient complex SA (CSA) mechanism. Benefiting from the conjugated symmetric property of DFT, CSA is capable to model the high-order contextual information with less than half computations of naive SA. To overcome the gradient explosion in Fourier complex field, we replace the Softmax function with the carefully designed Logmax function to normalize the attention map of CSA and stabilize the gradient propagation. By stacking various layers of CSA blocks, we propose the Fourier complex Transformer (FCT) model to learn global contextual information from VHR aerial images following the hierarchical manners. Universal experiments conducted on commonly used RS classification datasets demonstrate the effectiveness and efficiency of FCT, especially on VHR RS images. The source code of FCT will be available at https://github.com/Gao-xiyuan/FCT. Yan Zhang 0108, Xiyuan Gao, Qingyan Duan, Jiaxu Leng, Xiao Pu 0002, Xinbo Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | DecomFormer: Decompose Self-Attention Via Fourier Transform for VHR Aerial Image Scene ClassificationabstractVery high-resolution (VHR) aerial image scene classification is an essential task for aerial image understanding. Although transformer-based models have demonstrated strong ability in natural image classification, transformer-based methods on VHR aerial image tasks are still lack of concern because the complexity of self-attention in the transformer grows quadratically with the image resolution. To address this issue, we decompose the self-attention via Fourier Transform and propose a novel Fourier self-attention (FSA) mechanism. Based on FSA, we design a highly efficient network named DecomFormer, which learns contextual relationships in the real part and imaginary part of the Fourier field, respectively. Theoretically, the DecomFormer reduces the complexity of the naive self-attention mechanism from O(n2) to O(nlog(n)). Universal experiments on public VHR aerial image classification benchmarks demonstrated the DecomFormer’s efficiency, especially on images with very high-resolution. Yan Zhang 0108, Xiyuan Gao, Xiao Pu 0002, Xinbo Gao 0001 |
ICASSP | 2 |
| 2022 | Deep CNN-based Inductive Transfer Learning for Sarcasm Detection in SpeechabstractSarcasm is a frequently used linguistic device which is expressed in a multitude of ways, both with acoustic cues (including pitch, intonation, intensity, etc.) and visual cues (including facial expression, eye gaze, etc.). While cues used in the expression of sarcasm are well-described in the literature, there is a striking paucity of attempts to perform automatic sarcasm detection in speech. To explore this gap, we elaborate a methodology of implementing Inductive Transfer Learning (ITL) based on pre-trained Deep Convolutional Neural Networks (DCNNs) to detect sarcasm in speech. To those ends, the multimodal dataset MUStARD is used as a target dataset in this study. The two selected pre-trained DCNN models used are Xception and VGGish, which we trained on visual and audio datasets. Results show that VGGish, which is applied as a feature extractor in the experiment, performs better than Xception, which has its convolutional layers and pooling layers retrained. Both models achieve a higher F-score compared to the baseline Support Vector Machines (SVM) model by 7% and 5% in unimodal sarcasm detection in speech. Xiyuan Gao, Shekhar Nayak, Matt Coler |
INTERSPEECH | 1 |
| 2022 | DHT: Deformable Hybrid Transformer for Aerial Image SegmentationabstractDue to the strong ability to model global information, the transformer-based methods have shown remarkable improvements in image segmentation tasks. However, the self-attention mechanism in the transformer is computationally expensive and relies on pre-trained parameters. Moreover, the transformer method is weak in modeling local information, which is unfavorable for accurately segmenting objects from high-resolution aerial images. To this end, an efficient deformable orientational self-attention (DoA) is proposed to simultaneously extract the global information and the local information. Besides, for parameter efficiency, we design a depthwise channel self-attention (DcA) to model the contextual information among channels. Combining with the DoA and DcA, we propose the deformable hybrid transformer (DHT) to perform high-quality object segmentation on aerial images. Experiments on ISPRS Potsdam dataset and WHU building dataset illustrate that the proposed DHT can not only achieve state-of-the-art (SOTA) results but also markedly reduce the dependence of the transformer on pre-trained parameters. Yan Zhang 0108, Xiyuan Gao, Qingyan Duan, Lin Yuan 0002, Xinbo Gao 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2020 | Development of A Blockchain Framework for Virtual Clinical Trials
Yan Zhuang 0011, Lincoln Sheets, Xiyuan Gao, Zon-Yin Shae, Jeffrey J. P. Tsai, Chi-Ren Shyu |
AMIA | 3 |