VLDB 2026 Research / reviewers in the wild / expert
Yu-Ping Ruan
dblp:188/9011
· DBLP profile ↗
15ranked-venue papers
9as first author
10since 2021 · last 2025
0000-0002-9800-3271ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 6 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Ensuring Pre-Fusion Modality Consistency: A New Approach to Multimodal Sentiment DetectionabstractWith the growing diversity of data formats on social media, such as text, images, and videos, there is a growing need to analyze sentiment from multiple modalities. Multimodal sentiment detection, which aims to identify users’ sentiment by jointly modeling information from different modalities, has thus attracted increasing attention. However, most existing multimodal sentiment detection methods fuse multimodal information directly after the unimodal encoding and overlook the modality consistency of multimodal vector spaces before the fusion, which may damage the accuracy of multimodal sentiment detection. To address this issue, we propose a contrastive learning-based multimodal sentiment detection model termed EPMC which can map the representations of different modalities into a unified semantic space before fusion. EPMC operates in two stages, i.e., pre-training stage and fine-tuning stage. At the pre-training stage, we designed a cross-modal transformation module to map different modalities into a unified feature space. Meanwhile, to further capture the relationship between the cross-modal transformation vectors and the unimodal encoding vectors, we propose a multimodal consistency contrastive learning task that helps the model discern and amplify the cross-modal similarity between different modalities, thereby learning more discriminative features for sentiment detection. At the fine-tuning stage, EPMC is iteratively refined using the learned multimodal representation and guided by the cross-entropy loss. Extensive experiments conducted on three public multimodal datasets validate the effectiveness of EPMC model. The official implementation of EPMC is released at https://github.com/ADMIS-TONGJI/EPMC . Yulou Shu, Wengen Li, Yu-Ping Ruan, Wuchao Liu, Yichao Zhang 0001, Jihong Guan, Shuigeng Zhou |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2024 | RedCore: Relative Advantage Aware Cross-Modal Representation Learning for Missing Modalities with Imbalanced Missing RatesabstractMultimodal learning is susceptible to modality missing, which poses a major obstacle for its practical applications and, thus, invigorates increasing research interest. In this paper, we investigate two challenging problems: 1) when modality missing exists in the training data, how to exploit the incomplete samples while guaranteeing that they are properly supervised? 2) when the missing rates of different modalities vary, causing or exacerbating the imbalance among modalities, how to address the imbalance and ensure all modalities are well-trained. To tackle these two challenges, we first introduce the variational information bottleneck (VIB) method for the cross-modal representation learning of missing modalities, which capitalizes on the available modalities and the labels as supervision. Then, accounting for the imbalanced missing rates, we define relative advantage to quantify the advantage of each modality over others. Accordingly, a bi-level optimization problem is formulated to adaptively regulate the supervision of all modalities during training. As a whole, the proposed approach features Relative advantage aware Cross-modal representation learning (abbreviated as RedCore) for missing modalities with imbalanced missing rates. Extensive empirical results demonstrate that RedCore outperforms competing models in that it exhibits superior robustness against either large or imbalanced missing rates. The code is available at: https://github.com/sunjunaimer/RedCore. Shoukang Han, Yu-Ping Ruan, Taihao Li |
AAAI | 4 |
| 2024 | Fusing Modality-Specific Representations and Decisions for Multimodal Emotion RecognitionabstractMultimodal emotion recognition (MER) is important for building humanoid chatbots and has gained increasing attention in recent years. Existing studies have proven that extracting better modality-specific representations, which keep both commonality and individuality information of different modalities, is important for the MER task. However, all these works are restricted in making final predictions based on fusing modality-specific representations, and the effectiveness of the modality-specific decisions has not been studied. In this paper, we propose for the first time to fuse both the modality-specific representations and decisions for the MER task and design a bi-channel fusing network (BCFN). Specifically, a BCFN model first extracts and mixes the modality-specific representations and decisions in two convolutional blocks respectively, and then fuses the two joint multimodal features for the final decision. Extensive experiments are conducted on two MER benchmark datasets with state-of-the-art models employed as baselines for comparison. The empirical results demonstrate the superiority of our proposed BCFN model and confirm the effectiveness of incorporating modality-specific decisions for the MER task. Yu-Ping Ruan, Shoukang Han, Taihao Li |
ICASSP | 1 |
| 2024 | Multi-Modal Emotion Recognition Using Multiple Acoustic Features and Dual Cross-Modal TransformerabstractMulti-modal emotion recognition (MER) using speech and text has attracted extensive attention because of the easy availability of data for these two modalities. Recently, the self-surprised learning (SSL) pre-trained model has become the state-of-the-art (SOTA) method for the extraction of acoustic and textual features. However, the SSL speech representation may lose some important paralinguistic information, resulting in limited speech knowledge for MER. In this paper, we propose to adopt two kinds of acoustic features (i.e., the SSL representation and the spectral feature) as inputs to comprehensively extract speech characteristics. In addition, a dual cross-modal Transformer module is presented to model the interaction on the unaligned sequences between the textual feature and two acoustic features. Moreover, we introduce a blended loss including two uni-modal losses to better extract the uni-modal information. Experiments conducted on the widely used IEMOCAP dataset indicate that our proposed method achieves the SOTA performance compared with previous methods. Pengcheng Yue, Leyuan Qu, Taihao Li, Yu-Ping Ruan |
ICASSP | 5 |
| 2024 | Weakly Correlated Multimodal Sentiment Analysis: New Dataset and Topic-Oriented ModelabstractExisting multimodal sentiment analysis models focus more on fusing highly correlated image-text pairs, and thus achieves unsatisfactory performance on multimodal social media data which usually manifests weak correlations between different modalities. To address this issue, we first build a large multimodal social media sentiment analysis dataset RU-Senti which contains more than 100,000 image-text pairs with sentiment labels. Then, we proposed a topic-oriented model (TOM) which assumes that text is usually related to a certain portion of the image contents and significant variances exist in sentiment distribution across diverse topics. TOM learns the topic information from textual content and designs a topic-oriented feature alignment module to extract textual semantics correlated information from images, thus achieving the alignment between two modalities. Then, TOM utilizes a transformer encoder initialized with the parameters from a pre-trained vision-language model to fuse the multimodal features for sentiment prediction. According to the experiments over the public MVSA-Multiple dataset and our RU-Senti dataset, RU-Senti is of high suitability for studying weakly correlated multimodal sentiment analysis, and the proposed TOM model also largely outperforms the SOTA mulitimodal sentiment analysis methods and pre-trained vision-language models. Wuchao Liu, Wengen Li, Yu-Ping Ruan, Yulou Shu, Yina Li, Caili Yu, Yichao Zhang 0001, Jihong Guan, Shuigeng Zhou |
IEEE Trans. Affect. Comput. | 3 |
| 2023 | Layer-wise Fusion with Modality Independence Modeling for Multi-modal Emotion RecognitionabstractJun Sun, Shoukang Han, Yu-Ping Ruan, Xiaoning Zhang, Shu-Kai Zheng, Yulong Liu, Yuxin Huang, Taihao Li. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Shoukang Han, Yu-Ping Ruan, Shukai Zheng, Taihao Li |
ACL (1) | 3 |
| 2023 | Capsule Network with Label Dependency Modeling for Multi-Label Emotion ClassificationabstractThis paper proposes a simple-yet-efficient model, called capsule network with label dependency modeling (CapsLDM), for the task of multi-label emotion classification (MLEC) in text, in which multiple emotion categories can be assigned to the input data instance (e.g., a sentence). Unlike the traditional single-label emotion classification, the modeling of label (i.e., emotion) dependency plays an important role in MLEC, since the co-existing emotions in an utterance are not independent of each other. The capsule network has been successfully applied to many multi-label classification scenarios, however, the modeling of label dependency has not been considered in existing work. In our proposed CapsLDM model, we add similarity regularization terms on both the dynamic routing weights and the instance vectors of emotion capsules by exploiting the co-occurrence information of emotion labels, which resembles the dependency between different emotion categories for a certain input instance. Extensive experiments are conducted on four MLEC benchmark datasets with state-of-the-art models employed as baselines for comparison. The empirical results demonstrate the superiority of our proposed CapsLDM model and confirm the effectiveness of label dependency modeling in CapsLDM for the MLEC task. Yu-Ping Ruan, Taihao Li |
ECAI | 1 |
| 2023 | Emotion-Regularized Conditional Variational Autoencoder for Emotional Response GenerationabstractThis article presents an emotion-regularized conditional variational autoencoder (Emo-CVAE) model for generating emotional conversation responses. In conventional CVAE-based emotional response generation, emotion labels are simply used as additional conditions in prior, posterior and decoder networks. Considering that emotion styles are naturally entangled with semantic contents in the language space, the Emo-CVAE model utilizes emotion labels to regularize the CVAE latent space by introducing an extra emotion prediction network. In the training stage, the estimated latent variables are required to predict the emotion labels and token sequences of the input responses simultaneously. Experimental results show that our Emo-CVAE model can learn a more informative and structured latent space than a conventional CVAE model and output responses with better content and emotion performance than baseline CVAE and sequence-to-sequence (Seq2Seq) models. Yu-Ping Ruan, Zhen-Hua Ling |
IEEE Trans. Affect. Comput. | 1 |
| 2022 | Hierarchical and Multi-View Dependency Modelling Network for Conversational Emotion RecognitionabstractThis paper proposes a new model, called hierarchical and multi-view dependency modelling network (HMVDM), for the task of emotion recognition in conversations (ERC). The modelling of conversational context plays an important role in ERC, especially for the multi-turn and multi-speaker conversations which hold complex dependency between different speakers. In our proposed HMVDM1, we model the dependency between different speakers at both tokenlevel and utterance-level. Specifically, the HMVDM model has a hierarchical structure with two main modules: 1) token-level dependency modelling module (TDM), which aims to learn the long-range token-level dependency between different utterances in a speaker-aware manner and output the utterance representation; 2) utterance-level dependency modelling module (UDM), which accepts the utterance representation from TDM as inputs and aims to learn the utterance-level dependency from intra-, inter-, and global-speaker(s) view simultaneously. Extensive experiments are conducted on four ERC benchmark datasets with state-of-the-art models employed as baselines for comparison. The empirical results demonstrate the superiority of our proposed HMVDM model and confirm the importance of hierarchical and multi-view context dependency modelling for ERC. Yu-Ping Ruan, Shukai Zheng, Taihao Li, Guanxiong Pei |
ICASSP | 1 |
| 2021 | Deep Contextualized Utterance Representations for Response Selection and Dialogue AnalysisabstractThe NOESIS II challenge, as the Track 2 in the Eighth Dialogue System Technology Challenge (DSTC 8), is the extension of Track 1 in DSTC 7. Three new elements are incorporated into the extended track, i.e., dialogue with multiple participants, dialogue success, and dialogue disentanglement. These are vital for the creation of a deployed task-oriented dialogue system. This track is divided into four subtasks, the first two of which are evaluated in the form of response selection and the last two focus on dialogue analysis. This paper describes our methods developed for these four subtasks, which all employ deep contextualized utterance representations to make models aware of contextual information and to keep the intrinsic property of multi-turn dialogue systems. In the released evaluation results of Track 2 in DSTC 8, our proposed methods ranked fourth in subtask 1, third in subtask 2, and first in subtask 3 and subtask 4 respectively. In addition to the challenge tasks, we also compare our proposed methods with previous ones on public benchmark datasets. Experimental results show that our proposed methods outperform existing ones by large margins and achieve new state-of-the-art performances on multi-turn response selection and dialogue disentanglement. Jia-Chen Gu, Tianda Li, Zhen-Hua Ling, Quan Liu 0003, Zhiming Su, Yu-Ping Ruan, Xiaodan Zhu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2020 | Generating diverse conversation responses by creating and ranking multiple candidates
Yu-Ping Ruan, Zhen-Hua Ling, Xiaodan Zhu 0001, Quan Liu 0003, Jia-Chen Gu |
Comput. Speech Lang. | 1 |
| 2020 | Condition-Transforming Variational Autoencoder for Generating Diverse Short Text ConversationsabstractIn this article, conditional-transforming variational autoencoders (CTVAEs) are proposed for generating diverse short text conversations. In conditional variational autoencoders (CVAEs), the prior distribution of latent variable z follows a multivariate Gaussian distribution with mean and variance modulated by the input conditions. Previous work found that this distribution tended to become condition-independent in practical applications. Thus, this article designs CTVAEs to enhance the influence of conditions in CVAEs. In a CTVAE model, the latent variable z is sampled by performing a non-linear transformation on the combination of the input conditions and the samples from a condition-independent prior distribution N (0, I). In our experiments using a Chinese Sina Weibo dataset, the CTVAE model derives z samples for decoding with better condition-dependency than that of the CVAE model. The earth mover’s distance (EMD) between the distributions of the latent variable z at the training stage, and the testing stage is also reduced by using the CTVAE model. In subjective preference tests, our proposed CTVAE model performs significantly better than CVAE and sequence-to-sequence (Seq2Seq) models on generating diverse, informative, and topic-relevant responses. Yu-Ping Ruan, Zhen-Hua Ling, Xiaodan Zhu 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2019 | Condition-transforming Variational Autoencoder for Conversation Response GenerationabstractThis paper proposes a new model, called condition-transforming variational autoencoder (CTVAE), to improve the performance of conversation response generation using conditional variational autoencoders (CVAEs). In conventional CVAEs , the prior distribution of latent variable z follows a multivariate Gaussian distribution with mean and variance modulated by the input conditions. Previous work found that this distribution tends to become condition-independent in practical application. In our proposed CTVAE model, the latent variable z is sampled by performing a non-linear transformation on the combination of the input conditions and the samples from a condition-independent prior distribution N(0,I). In our objective evaluations, the CTVAE model outperforms the CVAE model on fluency metrics and surpasses a sequence-to-sequence (Seq2Seq) model on diversity metrics. In subjective preference tests, our proposed CTVAE model performs significantly better than CVAE and Seq2Seq models on generating fluency, informative and topic relevant responses. Yu-Ping Ruan, Zhen-Hua Ling, Quan Liu 0003, Zhigang Chen 0003, Nitin Indurkhya |
ICASSP | 1 |
| 2018 | A Sequential Neural Encoder With Latent Structured Description for Modeling SentencesabstractIn this paper, we propose a sequential neural encoder with latent structured description (SNELSD) for modeling sentences. This model introduces latent chunk-level representations into conventional sequential neural encoders, i.e., recurrent neural networks with long short-term memory (LSTM) units, to consider the compositionality of languages in semantic modeling. An SNELSD model has a hierarchical structure that includes a detection layer and a description layer. The detection layer predicts the boundaries of latent word chunks in an input sentence and derives a chunk-level vector for each word. The description layer utilizes modified LSTM units to process these chunk-level vectors in a recurrent manner and produces sequential encoding outputs. These output vectors are further concatenated with word vectors or the outputs of a chain LSTM encoder to obtain the final sentence representation. All the model parameters are learned in an end-to-end manner without a dependency on additional text chunking or syntax parsing. A natural language inference task and a sentiment analysis task are adopted to evaluate the performance of our proposed model. The experimental results demonstrate the effectiveness of the proposed SNELSD model on exploring task-dependent chunking patterns during the semantic modeling of sentences. Furthermore, the proposed method achieves better performance than conventional chain LSTMs and tree-structured LSTMs on both tasks. Yu-Ping Ruan, Qian Chen 0003, Zhen-Hua Ling |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2016 | Exploring Semantic Representation in Brain Activity Using Word EmbeddingsabstractIn this paper, we utilize distributed word representations (i.e., word embeddings) to analyse the representation of semantics in brain activity.The brain activity data were recorded using functional magnetic resonance imaging (fMRI) when subjects were viewing words.First, we analysed the functional selectivity of different cortex areas by calculating the correlations between neural responses and several types of word representations, including skipgram word embeddings, visual semantic vectors, and primary visual features.The results demonstrated consistency with existing neuroscientific knowledge.Second, we utilized behavioural data as the semantic ground truth to measure their relevance with brain activity.A method to estimate word embeddings under the constraints of brain activity similarities is further proposed based on the semantic word embedding (SWE) model.The experimental results show that the brain activity data are significantly correlated with the behavioural data of human judgements on semantic similarity.The correlations between the estimated word embeddings and the semantic ground truth can be effectively improved after integrating the brain activity data for learning, which implies that semantic patterns in neural representations may exist that have not been fully captured by state-of-the-art word embeddings derived from text corpora. Yu-Ping Ruan, Zhen-Hua Ling, Yu Hu 0003 |
EMNLP | 1 |