Yu-Ping Ruan

dblp:188/9011 · DBLP profile ↗
← Back
15ranked-venue papers
9as first author
10since 2021 · last 2025
0000-0002-9800-3271ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 6 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Ensuring Pre-Fusion Modality Consistency: A New Approach to Multimodal Sentiment Detection
abstract
With the growing diversity of data formats on social media, such as text, images, and videos, there is a growing need to analyze sentiment from multiple modalities. Multimodal sentiment detection, which aims to identify users’ sentiment by jointly modeling information from different modalities, has thus attracted increasing attention. However, most existing multimodal sentiment detection methods fuse multimodal information directly after the unimodal encoding and overlook the modality consistency of multimodal vector spaces before the fusion, which may damage the accuracy of multimodal sentiment detection. To address this issue, we propose a contrastive learning-based multimodal sentiment detection model termed EPMC which can map the representations of different modalities into a unified semantic space before fusion. EPMC operates in two stages, i.e., pre-training stage and fine-tuning stage. At the pre-training stage, we designed a cross-modal transformation module to map different modalities into a unified feature space. Meanwhile, to further capture the relationship between the cross-modal transformation vectors and the unimodal encoding vectors, we propose a multimodal consistency contrastive learning task that helps the model discern and amplify the cross-modal similarity between different modalities, thereby learning more discriminative features for sentiment detection. At the fine-tuning stage, EPMC is iteratively refined using the learned multimodal representation and guided by the cross-entropy loss. Extensive experiments conducted on three public multimodal datasets validate the effectiveness of EPMC model. The official implementation of EPMC is released at https://github.com/ADMIS-TONGJI/EPMC .
Yulou Shu, Wengen Li, Yu-Ping Ruan, Wuchao Liu, Yichao Zhang 0001, Jihong Guan, Shuigeng Zhou
ACM Trans. Intell. Syst. Technol.3
2024 RedCore: Relative Advantage Aware Cross-Modal Representation Learning for Missing Modalities with Imbalanced Missing Rates
abstract
Multimodal learning is susceptible to modality missing, which poses a major obstacle for its practical applications and, thus, invigorates increasing research interest. In this paper, we investigate two challenging problems: 1) when modality missing exists in the training data, how to exploit the incomplete samples while guaranteeing that they are properly supervised? 2) when the missing rates of different modalities vary, causing or exacerbating the imbalance among modalities, how to address the imbalance and ensure all modalities are well-trained. To tackle these two challenges, we first introduce the variational information bottleneck (VIB) method for the cross-modal representation learning of missing modalities, which capitalizes on the available modalities and the labels as supervision. Then, accounting for the imbalanced missing rates, we define relative advantage to quantify the advantage of each modality over others. Accordingly, a bi-level optimization problem is formulated to adaptively regulate the supervision of all modalities during training. As a whole, the proposed approach features Relative advantage aware Cross-modal representation learning (abbreviated as RedCore) for missing modalities with imbalanced missing rates. Extensive empirical results demonstrate that RedCore outperforms competing models in that it exhibits superior robustness against either large or imbalanced missing rates. The code is available at: https://github.com/sunjunaimer/RedCore.
Shoukang Han, Yu-Ping Ruan, Taihao Li
AAAI4
2024 Fusing Modality-Specific Representations and Decisions for Multimodal Emotion Recognition
abstract
Multimodal emotion recognition (MER) is important for building humanoid chatbots and has gained increasing attention in recent years. Existing studies have proven that extracting better modality-specific representations, which keep both commonality and individuality information of different modalities, is important for the MER task. However, all these works are restricted in making final predictions based on fusing modality-specific representations, and the effectiveness of the modality-specific decisions has not been studied. In this paper, we propose for the first time to fuse both the modality-specific representations and decisions for the MER task and design a bi-channel fusing network (BCFN). Specifically, a BCFN model first extracts and mixes the modality-specific representations and decisions in two convolutional blocks respectively, and then fuses the two joint multimodal features for the final decision. Extensive experiments are conducted on two MER benchmark datasets with state-of-the-art models employed as baselines for comparison. The empirical results demonstrate the superiority of our proposed BCFN model and confirm the effectiveness of incorporating modality-specific decisions for the MER task.
Yu-Ping Ruan, Shoukang Han, Taihao Li
ICASSP1
2024 Multi-Modal Emotion Recognition Using Multiple Acoustic Features and Dual Cross-Modal Transformer
abstract
Multi-modal emotion recognition (MER) using speech and text has attracted extensive attention because of the easy availability of data for these two modalities. Recently, the self-surprised learning (SSL) pre-trained model has become the state-of-the-art (SOTA) method for the extraction of acoustic and textual features. However, the SSL speech representation may lose some important paralinguistic information, resulting in limited speech knowledge for MER. In this paper, we propose to adopt two kinds of acoustic features (i.e., the SSL representation and the spectral feature) as inputs to comprehensively extract speech characteristics. In addition, a dual cross-modal Transformer module is presented to model the interaction on the unaligned sequences between the textual feature and two acoustic features. Moreover, we introduce a blended loss including two uni-modal losses to better extract the uni-modal information. Experiments conducted on the widely used IEMOCAP dataset indicate that our proposed method achieves the SOTA performance compared with previous methods.
Pengcheng Yue, Leyuan Qu, Taihao Li, Yu-Ping Ruan
ICASSP5
2024 Weakly Correlated Multimodal Sentiment Analysis: New Dataset and Topic-Oriented Model
abstract
Existing multimodal sentiment analysis models focus more on fusing highly correlated image-text pairs, and thus achieves unsatisfactory performance on multimodal social media data which usually manifests weak correlations between different modalities. To address this issue, we first build a large multimodal social media sentiment analysis dataset RU-Senti which contains more than 100,000 image-text pairs with sentiment labels. Then, we proposed a topic-oriented model (TOM) which assumes that text is usually related to a certain portion of the image contents and significant variances exist in sentiment distribution across diverse topics. TOM learns the topic information from textual content and designs a topic-oriented feature alignment module to extract textual semantics correlated information from images, thus achieving the alignment between two modalities. Then, TOM utilizes a transformer encoder initialized with the parameters from a pre-trained vision-language model to fuse the multimodal features for sentiment prediction. According to the experiments over the public MVSA-Multiple dataset and our RU-Senti dataset, RU-Senti is of high suitability for studying weakly correlated multimodal sentiment analysis, and the proposed TOM model also largely outperforms the SOTA mulitimodal sentiment analysis methods and pre-trained vision-language models.
Wuchao Liu, Wengen Li, Yu-Ping Ruan, Yulou Shu, Yina Li, Caili Yu, Yichao Zhang 0001, Jihong Guan, Shuigeng Zhou
IEEE Trans. Affect. Comput.3
2023 Layer-wise Fusion with Modality Independence Modeling for Multi-modal Emotion Recognition
abstract
Jun Sun, Shoukang Han, Yu-Ping Ruan, Xiaoning Zhang, Shu-Kai Zheng, Yulong Liu, Yuxin Huang, Taihao Li. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Shoukang Han, Yu-Ping Ruan, Shukai Zheng, Taihao Li
ACL (1)3
2023 Capsule Network with Label Dependency Modeling for Multi-Label Emotion Classification
abstract
This paper proposes a simple-yet-efficient model, called capsule network with label dependency modeling (CapsLDM), for the task of multi-label emotion classification (MLEC) in text, in which multiple emotion categories can be assigned to the input data instance (e.g., a sentence). Unlike the traditional single-label emotion classification, the modeling of label (i.e., emotion) dependency plays an important role in MLEC, since the co-existing emotions in an utterance are not independent of each other. The capsule network has been successfully applied to many multi-label classification scenarios, however, the modeling of label dependency has not been considered in existing work. In our proposed CapsLDM model, we add similarity regularization terms on both the dynamic routing weights and the instance vectors of emotion capsules by exploiting the co-occurrence information of emotion labels, which resembles the dependency between different emotion categories for a certain input instance. Extensive experiments are conducted on four MLEC benchmark datasets with state-of-the-art models employed as baselines for comparison. The empirical results demonstrate the superiority of our proposed CapsLDM model and confirm the effectiveness of label dependency modeling in CapsLDM for the MLEC task.
Yu-Ping Ruan, Taihao Li
ECAI1
2023 Emotion-Regularized Conditional Variational Autoencoder for Emotional Response Generation
abstract
This article presents an emotion-regularized conditional variational autoencoder (Emo-CVAE) model for generating emotional conversation responses. In conventional CVAE-based emotional response generation, emotion labels are simply used as additional conditions in prior, posterior and decoder networks. Considering that emotion styles are naturally entangled with semantic contents in the language space, the Emo-CVAE model utilizes emotion labels to regularize the CVAE latent space by introducing an extra emotion prediction network. In the training stage, the estimated latent variables are required to predict the emotion labels and token sequences of the input responses simultaneously. Experimental results show that our Emo-CVAE model can learn a more informative and structured latent space than a conventional CVAE model and output responses with better content and emotion performance than baseline CVAE and sequence-to-sequence (Seq2Seq) models.
Yu-Ping Ruan, Zhen-Hua Ling
IEEE Trans. Affect. Comput.1
2022 Hierarchical and Multi-View Dependency Modelling Network for Conversational Emotion Recognition
abstract
This paper proposes a new model, called hierarchical and multi-view dependency modelling network (HMVDM), for the task of emotion recognition in conversations (ERC). The modelling of conversational context plays an important role in ERC, especially for the multi-turn and multi-speaker conversations which hold complex dependency between different speakers. In our proposed HMVDM1, we model the dependency between different speakers at both tokenlevel and utterance-level. Specifically, the HMVDM model has a hierarchical structure with two main modules: 1) token-level dependency modelling module (TDM), which aims to learn the long-range token-level dependency between different utterances in a speaker-aware manner and output the utterance representation; 2) utterance-level dependency modelling module (UDM), which accepts the utterance representation from TDM as inputs and aims to learn the utterance-level dependency from intra-, inter-, and global-speaker(s) view simultaneously. Extensive experiments are conducted on four ERC benchmark datasets with state-of-the-art models employed as baselines for comparison. The empirical results demonstrate the superiority of our proposed HMVDM model and confirm the importance of hierarchical and multi-view context dependency modelling for ERC.
Yu-Ping Ruan, Shukai Zheng, Taihao Li, Guanxiong Pei
ICASSP1
2021 Deep Contextualized Utterance Representations for Response Selection and Dialogue Analysis
abstract
The NOESIS II challenge, as the Track 2 in the Eighth Dialogue System Technology Challenge (DSTC 8), is the extension of Track 1 in DSTC 7. Three new elements are incorporated into the extended track, i.e., dialogue with multiple participants, dialogue success, and dialogue disentanglement. These are vital for the creation of a deployed task-oriented dialogue system. This track is divided into four subtasks, the first two of which are evaluated in the form of response selection and the last two focus on dialogue analysis. This paper describes our methods developed for these four subtasks, which all employ deep contextualized utterance representations to make models aware of contextual information and to keep the intrinsic property of multi-turn dialogue systems. In the released evaluation results of Track 2 in DSTC 8, our proposed methods ranked fourth in subtask 1, third in subtask 2, and first in subtask 3 and subtask 4 respectively. In addition to the challenge tasks, we also compare our proposed methods with previous ones on public benchmark datasets. Experimental results show that our proposed methods outperform existing ones by large margins and achieve new state-of-the-art performances on multi-turn response selection and dialogue disentanglement.
Jia-Chen Gu, Tianda Li, Zhen-Hua Ling, Quan Liu 0003, Zhiming Su, Yu-Ping Ruan, Xiaodan Zhu 0001
IEEE ACM Trans. Audio Speech Lang. Process.6
2020 Generating diverse conversation responses by creating and ranking multiple candidates
Yu-Ping Ruan, Zhen-Hua Ling, Xiaodan Zhu 0001, Quan Liu 0003, Jia-Chen Gu
Comput. Speech Lang.1
2020 Condition-Transforming Variational Autoencoder for Generating Diverse Short Text Conversations
abstract
In this article, conditional-transforming variational autoencoders (CTVAEs) are proposed for generating diverse short text conversations. In conditional variational autoencoders (CVAEs), the prior distribution of latent variable z follows a multivariate Gaussian distribution with mean and variance modulated by the input conditions. Previous work found that this distribution tended to become condition-independent in practical applications. Thus, this article designs CTVAEs to enhance the influence of conditions in CVAEs. In a CTVAE model, the latent variable z is sampled by performing a non-linear transformation on the combination of the input conditions and the samples from a condition-independent prior distribution N (0, I). In our experiments using a Chinese Sina Weibo dataset, the CTVAE model derives z samples for decoding with better condition-dependency than that of the CVAE model. The earth mover’s distance (EMD) between the distributions of the latent variable z at the training stage, and the testing stage is also reduced by using the CTVAE model. In subjective preference tests, our proposed CTVAE model performs significantly better than CVAE and sequence-to-sequence (Seq2Seq) models on generating diverse, informative, and topic-relevant responses.
Yu-Ping Ruan, Zhen-Hua Ling, Xiaodan Zhu 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2019 Condition-transforming Variational Autoencoder for Conversation Response Generation
abstract
This paper proposes a new model, called condition-transforming variational autoencoder (CTVAE), to improve the performance of conversation response generation using conditional variational autoencoders (CVAEs). In conventional CVAEs , the prior distribution of latent variable z follows a multivariate Gaussian distribution with mean and variance modulated by the input conditions. Previous work found that this distribution tends to become condition-independent in practical application. In our proposed CTVAE model, the latent variable z is sampled by performing a non-linear transformation on the combination of the input conditions and the samples from a condition-independent prior distribution N(0,I). In our objective evaluations, the CTVAE model outperforms the CVAE model on fluency metrics and surpasses a sequence-to-sequence (Seq2Seq) model on diversity metrics. In subjective preference tests, our proposed CTVAE model performs significantly better than CVAE and Seq2Seq models on generating fluency, informative and topic relevant responses.
Yu-Ping Ruan, Zhen-Hua Ling, Quan Liu 0003, Zhigang Chen 0003, Nitin Indurkhya
ICASSP1
2018 A Sequential Neural Encoder With Latent Structured Description for Modeling Sentences
abstract
In this paper, we propose a sequential neural encoder with latent structured description (SNELSD) for modeling sentences. This model introduces latent chunk-level representations into conventional sequential neural encoders, i.e., recurrent neural networks with long short-term memory (LSTM) units, to consider the compositionality of languages in semantic modeling. An SNELSD model has a hierarchical structure that includes a detection layer and a description layer. The detection layer predicts the boundaries of latent word chunks in an input sentence and derives a chunk-level vector for each word. The description layer utilizes modified LSTM units to process these chunk-level vectors in a recurrent manner and produces sequential encoding outputs. These output vectors are further concatenated with word vectors or the outputs of a chain LSTM encoder to obtain the final sentence representation. All the model parameters are learned in an end-to-end manner without a dependency on additional text chunking or syntax parsing. A natural language inference task and a sentiment analysis task are adopted to evaluate the performance of our proposed model. The experimental results demonstrate the effectiveness of the proposed SNELSD model on exploring task-dependent chunking patterns during the semantic modeling of sentences. Furthermore, the proposed method achieves better performance than conventional chain LSTMs and tree-structured LSTMs on both tasks.
Yu-Ping Ruan, Qian Chen 0003, Zhen-Hua Ling
IEEE ACM Trans. Audio Speech Lang. Process.1
2016 Exploring Semantic Representation in Brain Activity Using Word Embeddings
abstract
In this paper, we utilize distributed word representations (i.e., word embeddings) to analyse the representation of semantics in brain activity.The brain activity data were recorded using functional magnetic resonance imaging (fMRI) when subjects were viewing words.First, we analysed the functional selectivity of different cortex areas by calculating the correlations between neural responses and several types of word representations, including skipgram word embeddings, visual semantic vectors, and primary visual features.The results demonstrated consistency with existing neuroscientific knowledge.Second, we utilized behavioural data as the semantic ground truth to measure their relevance with brain activity.A method to estimate word embeddings under the constraints of brain activity similarities is further proposed based on the semantic word embedding (SWE) model.The experimental results show that the brain activity data are significantly correlated with the behavioural data of human judgements on semantic similarity.The correlations between the estimated word embeddings and the semantic ground truth can be effectively improved after integrating the brain activity data for learning, which implies that semantic patterns in neural representations may exist that have not been fully captured by state-of-the-art word embeddings derived from text corpora.
Yu-Ping Ruan, Zhen-Hua Ling, Yu Hu 0003
EMNLP1