VLDB 2026 Research / reviewers in the wild / expert
Sijie Mai
dblp:245/8717
· DBLP profile ↗
36ranked-venue papers
18as first author
32since 2021 · last 2026
0000-0001-9763-375XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 12 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 7 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Cosine Similarity: Magnitude-Aware CLIP for No-Reference Image Quality AssessmentabstractRecent efforts have repurposed the Contrastive Language-Image Pre-training (CLIP) model for No-Reference Image Quality Assessment (NR-IQA) by measuring the cosine similarity between the image embedding and textual prompts such as "a good photo" or "a bad photo." However, this semantic similarity overlooks a critical yet underexplored cue: the magnitude of the CLIP image features, which we empirically find to exhibit a strong correlation with perceptual quality. In this work, we introduce a novel adaptive fusion framework that complements cosine similarity with a magnitude-aware quality cue. Specifically, we first extract the absolute CLIP image features and apply a Box-Cox transformation to statistically normalize the feature distribution and mitigate semantic sensitivity. The resulting scalar summary serves as a semantically-normalized auxiliary cue that complements cosine-based prompt matching. To integrate both cues effectively, we further design a confidence-guided fusion scheme that adaptively weighs each term according to its relative strength. Extensive experiments on multiple benchmark IQA datasets demonstrate that our method consistently outperforms standard CLIP-based IQA and state-of-the-art baselines, without any task-specific training. Zhicheng Liao, Dongxu Wu, Zhenshan Shi, Sijie Mai, Hanwei Zhu, Lingyu Zhu 0006, Yuncheng Jiang 0004, Baoliang Chen |
AAAI | 4 |
| 2026 | Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference PerspectiveabstractMultimodal affective computing aims to predict humans' sentiment, emotion, intention, and opinion using language, acoustic, and visual modalities.However, current models often learn spurious correlations that harm generalization under distribution shifts or noisy modalities.To address this, we propose a causal modality-invariant representation (CmIR) learning framework for robust multimodal learning.At its core, we introduce a theoretically grounded disentanglement method that separates each modality into 'causal invariant representation' and 'environment-specific spurious representation' from a causal inference perspective.CmIR ensures that the learned invariant representations retain stable predictive relationships with labels across different environments while preserving sufficient information from the raw inputs via invariance constraint, mutual information constraint, and reconstruction constraint.Experiments across multiple multimodal benchmarks demonstrate that CmIR achieves state-of-theart performance.CmIR particularly excels on out-of-distribution data and noisy data, confirming its robustness and generalizability. Sijie Mai, Shiqin Han |
ACL (1) | 1 |
| 2026 | Tri-Modal grouping fusion network with optimal transport learning for unaligned multimodal sentiment analysis
Aolin Xiong, Sijie Mai, Haifeng Hu 0001 |
Expert Syst. Appl. | 2 |
| 2026 | Affection-Guided Bottleneck Diffusion for Missing Modality Issue in Multimodal Affective ComputingabstractMissing modality issue in multimodal affective computing severely hinders the robustness and performance of multimodal learning, particularly in real-world scenarios. Existing methods often fail in fully exploiting the remaining modalities, leading to noisy reconstruction process for the missing modalities and yielding suboptimal results. Besides, most of these methods rely on designing sophisticated networks to handle various missing scenarios, which prevents them from taking advantage of the original pre-trained multimodal networks trained for complete multimodal inputs. To address these challenges, we propose Affection-guided Bottleneck Diffusion (ABDiff), a novel approach leveraging score-based diffusion generative encoders to reconstruct missing modalities in the latent space without modification to the pre-trained fusion models. By incorporating self- and cross-attention mechanisms inside and among the missing and remaining modalities, ABDiff captures both modality-specific dynamics and cross-modal interactions during generation. Furthermore, an affection-guided information bottleneck is introduced to filter task-unrelated noise and modality-specific redundancy, stabilizing the generation process of missing modalities. The generated representations are seamlessly integrated with the remaining modalities into the pre-trained fusion networks. Extensive experiments on four public multimodal affective computing datasets demonstrate that ABDiff surpasses previous methods under both complete and incomplete modality scenarios. The code is released inhttps://github.com/RH-Lin/ABDiff. Ronghao Lin, Qiaolin He, Sijie Mai, Haifeng Hu 0001 |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | Supervised Attention Mechanism for Low-quality Multimodal DataabstractIn practical applications, multimodal data are often of low quality, with noisy modalities and missing modalities being typical forms that severely hinder model performance, robustness, and applicability.However, current studies address these issues separately.To this end, we propose a framework for multimodal affective computing that jointly addresses missing and noisy modalities to enhance model robustness in low-quality data scenarios.Specifically, we view missing modality as a special case of noisy modality, and propose a supervised attention framework.In contrast to traditional attention mechanisms that rely on main task loss to update the parameters, we design supervisory signals for the learning of attention weights, ensuring that attention mechanisms can focus on discriminative information and suppress noisy information.We further propose a ranking-based optimization strategy to compare the relative importance of different interactions by adding a ranking constraint for attention weights, avoiding training noise caused by inaccurate absolute labels.The proposed model consistently outperforms state-of-the-art baselines on multiple datasets under the settings of complete modalities, missing modalities, and noisy modalities. Sijie Mai, Shiqin Han, Haifeng Hu 0001 |
EMNLP | 1 |
| 2025 | Towards Explainable Fusion and Balanced Learning in Multimodal Sentiment AnalysisabstractMultimodal Sentiment Analysis (MSA) faces two critical challenges: the lack of interpretability in the decision logic of multimodal fusion and modality imbalance caused by disparities in inter-modal information density. To address these issues, we propose KAN-MCP, a novel framework that integrates the interpretability of Kolmogorov-Arnold Networks (KAN) with the robustness of the Multimodal Clean Pareto (MCPareto) framework. First, KAN leverages its univariate function decomposition to achieve transparent analysis of cross-modal interactions. This structural design allows direct inspection of feature transformations without relying on external interpretation tools, thereby ensuring both high expressiveness and interpretability. Second, the proposed MCPareto enhances robustness by addressing modality imbalance and noise interference. Specifically, we introduce the Dimensionality Reduction and Denoising Modal Information Bottleneck (DRD-MIB) method, which jointly denoises and reduces feature dimensionality. This approach provides KAN with discriminative low-dimensional inputs to reduce the modeling complexity of KAN while preserving critical sentiment-related information. Furthermore, MCPareto dynamically balances gradient contributions across modalities using the purified features output by DRD-MIB, ensuring lossless transmission of auxiliary signals and effectively alleviating modality imbalance. This synergy of interpretability and robustness not only achieves superior performance on benchmark datasets such as CMU-MOSI, CMU-MOSEI, and CH-SIMS v2 but also offers an intuitive visualization interface through KAN's interpretable architecture. Our code is released on https://github.com/LuoMSen/KAN-MCP. Miaosen Luo, Yuncheng Jiang 0004, Sijie Mai |
ACM Multimedia | 3 |
| 2025 | CyIN: Cyclic Informative Latent Space for Bridging Complete and Incomplete Multimodal LearningabstractMultimodal machine learning, mimicking the human brain’s ability to integrate various modalities has seen rapid growth. Most previous multimodal models are trained on perfectly paired multimodal input to reach optimal performance. In real‑world deployments, however, the presence of modality is highly variable and unpredictable, causing the pre-trained models in suffering significant performance drops and fail to remain robust with dynamic missing modalities circumstances. In this paper, we present a novel Cyclic INformative Learning framework (CyIN) to bridge the gap between complete and incomplete multimodal learning. Specifically, we firstly build an informative latent space by adopting token- and label-level Information Bottleneck (IB) cyclically among various modalities. Capturing task-related features with variational approximation, the informative bottleneck latents are purified for more efficient cross-modal interaction and multimodal fusion. Moreover, to supplement the missing information caused by incomplete multimodal input, we propose cross-modal cyclic translation by reconstruct the missing modalities with the remained ones through forward and reverse propagation process. With the help of the extracted and reconstructed informative latents, CyIN succeeds in jointly optimizing complete and incomplete multimodal learning in one unified model. Extensive experiments on 4 multimodal datasets demonstrate the superior performance of our method in both complete and diverse incomplete scenarios. Ronghao Lin, Qiaolin He, Sijie Mai, Aolin Xiong, Yap-Peng Tan, Haifeng Hu 0001 |
NeurIPS | 3 |
| 2025 | Learning by Comparing: Boosting Multimodal Affective Computing through Ordinal LearningabstractPrevious studies on multimodal affective computing primarily focus on approximating predictions to annotated labels, often neglecting the ordinal nature of affective states. In this paper, we address this issue by exploring ordinal learning, and a Multimodal Ordinal Affective Computing (MOAC) framework is designed to enhance the understanding of the nature of affective concepts. Specifically, we propose coarse-grained label-level ordinal learning that prompts the model to learn to compare in the label space, encouraging higher predictive values for samples annotated with larger labels over those with smaller labels. Moreover, a regularization loss is proposed to prevent the output distributions from deviating significantly from the annotated label distributions. Fine-grained feature-level ordinal learning is then performed via the feature difference operation and the neutral embedding. The former compares samples in the feature space, calculating the difference between features of different samples to generate 'new' features for a more robust training. The latter seeks to reduce the difficulty of prediction by estimating the difference between the target multimodal representations and a neutral reference. We first demonstrate MOAC in multimodal sentiment analysis, which is a regression task that aligns well with the function of ordinal learning. Then we extend MOAC to classification tasks including multimodal humor detection and sarcasm detection to evaluate its generalizability. Experiments suggest that MOAC outperforms state-of-the-art methods. Sijie Mai, Haifeng Hu 0001 |
WWW | 1 |
| 2025 | Injecting Multimodal Information Into Pre-Trained Language Model for Multimodal Sentiment AnalysisabstractWith the increasing availability of computational and data resources, numerous powerful pre-trained language models (PLMs) have emerged for natural language processing tasks. However, how to inject nonverbal modalities into PLMs to handle multimodal information remains a practical problem. In this paper, we explore the application of PLM on multimodal sentiment analysis from a different perspective. Unlike many recent methods that develop multimodal fusion layers that are sequential to attention layers, we investigate the effectiveness of cross-modal additive attention that is parallel to attention layers, which takes the language modality as dominant modality. Moreover, we devise a gating mechanism to control the flow of nonverbal information by estimating its discriminative level. In this way, we can prevent noisy multimodal information from damaging the performance of pre-trained language model. In our framework, nonverbal modalities serve as auxiliary roles to provide the model with additional information and improve the understanding of multimodal human language. Additionally, cross-modal margin and matching losses are proposed to align the distributions of various modalities and simultaneously retain modality-specific information, which to some extent address the shortcoming of contrastive learning loss. Comprehensive experiments show that our approach surpasses existing state-of-the-art methods on multimodal sentiment analysis and emotion recognition tasks. Sijie Mai, Aolin Xiong, Haifeng Hu 0001 |
IEEE Trans. Affect. Comput. | 1 |
| 2024 | Relation-dependent contrastive learning with cluster sampling for inductive relation prediction
Aolin Xiong, Sijie Mai, Haifeng Hu 0001 |
Neurocomputing | 3 |
| 2024 | Multimodal Boosting: Addressing Noisy Modalities and Identifying Modality ContributionabstractIn multimodal representation learning, different modalities do not contribute equally. Especially when learning with noisy modalities that convey non-discriminative information, the prediction based on multimodal representation is often biased and even ignores the knowledge from informative modalities. In this paper, we aim to address the noisy modality problem and balance the contributions of multiple modalities dynamically in a parallel format. Specifically, we construct multiple base learners and formulate our framework as a boosting-like algorithm, where different base learners focus on different aspects of multimodal learning. To identify the contributions of individual base learners, we develop a contribution learning network that dynamically determines the contribution and noise level of each base learner. In contrast to the commonly considered attention mechanism, we define the transformation of predictive loss as the supervision signal to train the contribution learning network, which enables more accurate learning of modality importance. We derive the final prediction by incorporating the predictions of base learners based on their contributions. Notably, different from late fusion, we devise a multimodal base learner to explore the cross-modal interactions. To update the network, we design the ‘complementary update mechanism’, where for each base learner, we assign higher weights to those samples that are incorrectly predicted by other base learners. In this way, we can leverage the available information to correctly predict each sample to the utmost extent and enable different base learners to learn different aspects of multimodal information. Extensive experiments demonstrate that the proposed method achieves superior performance on multimodal sentiment analysis and emotion recognition. Sijie Mai, Aolin Xiong, Haifeng Hu 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | Multimodal Reaction: Information Modulation for Cross-Modal Representation LearningabstractIn multimodal machine learning, proper handling of cross-modal information is essential for obtaining an ideal joint embedding. Despite the progress made by recent fusion strategies, we hold that before the fusion stage, the unimodal representation inevitably contains noise that may hinder the correct learning of cross-modal dynamics and affect multimodal fusion. It is worthwhile to investigate how the information is being utilized and how to make the full use of it. Rethinking the process of leveraging multiple modalities for the joint embedding, multimodal learning can be regarded as achemical reactionprocess and two steps may benefit learning: 1) purification to filter impurity, and 2) catalyst to facilitate learning. In this paper, we propose aMultimodalInformationModulation (MIM) learning framework to modulate the contribution and utilization of the cross-modal information, which identifies and handles the ‘impurity’ and ‘catalyst’ in multimodal learning. Specifically, a Unimodal Purification Network (UPN) is proposed to identify and explicitly filter out the impurity within each modality before fusion, which reduces the possibility of learning incorrect cross-modal dynamics. Besides, based on the intuition that useful information has the potential in the guidance of model updating, it plays a role to facilitate learning, which is achieved by the design of the Knowledge Guidance Scheme (KGS) considering both the intra- and inter-modal scenarios. Different to a majority of works that emphasize the role of useful information in the fusion and inference stage, KGS considers its potential role in assisting the representation learning of weaker components. Besides, it fully considers the modality dominance problem and sample variations for optimization. In short, MIM manages to modulate the useless/useful information to minimize/emphasize their contribution. Experimental results verify the effectiveness of the proposed method. Sijie Mai, Haifeng Hu 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | IIB-MIL: Integrated Instance-Level and Bag-Level Multiple Instances Learning with Label Disambiguation for Pathological Image Analysis
Qin Ren 0001, Yu Zhao 0009, Bingzhe Wu, Sijie Mai, Yueshan Huang, Yonghong He, Junzhou Huang, Jianhua Yao 0001 |
MICCAI (6) | 5 |
| 2023 | SC-AIR-BERT: a pre-trained single-cell model for predicting the antigen-binding specificity of the adaptive immune receptorabstractAccurately predicting the antigen-binding specificity of adaptive immune receptors (AIRs), such as T-cell receptors (TCRs) and B-cell receptors (BCRs), is essential for discovering new immune therapies. However, the diversity of AIR chain sequences limits the accuracy of current prediction methods. This study introduces SC-AIR-BERT, a pre-trained model that learns comprehensive sequence representations of paired AIR chains to improve binding specificity prediction. SC-AIR-BERT first learns the 'language' of AIR sequences through self-supervised pre-training on a large cohort of paired AIR chains from multiple single-cell resources. The model is then fine-tuned with a multilayer perceptron head for binding specificity prediction, employing the K-mer strategy to enhance sequence representation learning. Extensive experiments demonstrate the superior AUC performance of SC-AIR-BERT compared with current methods for TCR- and BCR-binding specificity prediction. Yu Zhao 0009, Xiaona Su, Sijie Mai, Chenchen Qin, Rongshan Yu, Jianhua Yao 0001 |
Briefings Bioinform. | 4 |
| 2023 | Hybrid Contrastive Learning of Tri-Modal Representation for Multimodal Sentiment AnalysisabstractThe wide application of smart devices enables the availability of multimodal data, which can be utilized in many tasks. In the field of multimodal sentiment analysis, most previous works focus on exploring intra- and inter-modal interactions. However, training a network with cross-modal information (language, audio and visual) is still challenging due to the modality gap. Besides, while learning dynamics within each sample draws great attention, the learning of inter-sample and inter-class relationships is neglected. Moreover, the size of datasets limits the generalization ability of the models. To address the afore-mentioned issues, we propose a novel framework HyCon for hybrid contrastive learning of tri-modal representation. Specifically, we simultaneously perform intra-/inter-modal contrastive learning and semi-contrastive learning, with which the model can fully explore cross-modal interactions, learn inter-sample and inter-class relationships, and reduce the modality gap. Besides, refinement term and modality margin are introduced to enable a better learning of unimodal pairs. Moreover, we devise pair selection mechanism to identify and assign weights to the informative negative and positive pairs. HyCon can naturally generate many training pairs for better generalization and reduce the negative effect of limited datasets. Extensive experiments demonstrate that our method outperforms baselines on multimodal sentiment analysis and emotion recognition. Sijie Mai, Shuangjia Zheng, Haifeng Hu 0001 |
IEEE Trans. Affect. Comput. | 1 |
| 2023 | Learning to Learn Better Unimodal Representations via Adaptive Multimodal Meta-LearningabstractMultimodal sentiment analysis is an emerging field of artificial intelligence. The most predominant approaches have made notable progress by designing sophisticated fusion architectures, exploring inter-modal interactions between modalities. However, these works tend to utilize a uniform optimization strategy for each modality, so that only sub-optimal unimodal representations are obtained for multimodal fusion. To address this issue, we propose a novel meta-learning based paradigm that can retain the advantages of unimodal existence and further boost the performance of multimodal fusion. Specifically, we introduce the Adaptive Multimodal Meta-Learning (AMML) to meta-learn the unimodal networks and adapt them for multimodal inference. AMML can (1) effectively obtain more optimized unimodal representation via meta-training on unimodal tasks, which adaptively adjusts the learning rate and assigns a more specific optimization procedure for each modality; (2) and adapt the optimized unimodal representations for multimodal fusion via meta-testing on multimodal tasks. Considering multimodal fusion often suffers from the distributional mismatches between features of different modalities due to heterogeneous nature of the signals, we implement a distribution transformation layer on unimodal representations to regularize the unimodal distributions. In this way, distribution gaps can be reduced to achieve a better effect of fusion. Extensive experiments on two widely-used datasets demonstrate that AMML achieves state-of-the-art performance. Sijie Mai, Haifeng Hu 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | Subgraph-Aware Few-Shot Inductive Link Prediction Via Meta-LearningabstractLink prediction for knowledge graphs aims to predict missing connections between entities. Prevailing methods are limited to a transductive setting and hard to process unseen entities. The recently proposed subgraph-based models provide alternatives to predict links from the subgraph structure surrounding a candidate triplet. However, these methods require abundant known facts of training triplets and perform poorly on relationships that only have a few triplets. In this paper, we propose Meta-iKG, a novel subgraph-based meta-learner for few-shot inductive relation reasoning. Meta-iKG utilizes local subgraphs to transfer subgraph-specific information and to rapidly learn transferable patterns via meta-gradients. In this way, we find the model can quickly adapt to few-shot relationships using only a handful of known facts with inductive settings. Moreover, we introduce a large-shot relation updating procedure to ensure that our model can generalize well to both few-shot and large-shot relations. We evaluate Meta-iKG on inductive benchmarks sampled from the NELL and Freebase, and the results show that Meta-iKG outperforms the currently state-of-the-art methods in both few-shot scenarios and standard inductive settings. Shuangjia Zheng, Sijie Mai, Haifeng Hu 0001, Yuedong Yang |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Multimodal Information Bottleneck: Learning Minimal Sufficient Unimodal and Multimodal RepresentationsabstractLearning effective joint embedding for cross-modal data has always been a focus in the field of multimodal machine learning. We argue that during multimodal fusion, the generated multimodal embedding may be redundant, and the discriminative unimodal information may be ignored, which often interferes with accurate prediction and leads to a higher risk of overfitting. Moreover, unimodal representations also contain noisy information that negatively influences the learning of cross-modal dynamics. To this end, we introduce the multimodal information bottleneck (MIB), aiming to learn a powerful and sufficient multimodal representation that is free of redundancy and to filter out noisy information in unimodal representations. Specifically, inheriting from the general information bottleneck (IB), MIB aims to learn the minimal sufficient representation for a given task by maximizing the mutual information between the representation and the target and simultaneously constraining the mutual information between the representation and the input data. Different from general IB, our MIB regularizes both the multimodal and unimodal representations, which is a comprehensive and flexible framework that is compatible with any fusion methods. We develop three MIB variants, namely, early-fusion MIB, late-fusion MIB, and complete MIB, to focus on different perspectives of information constraints. Experimental results suggest that the proposed method reaches state-of-the-art performance on the tasks of multimodal sentiment analysis and multimodal emotion recognition across three widely used datasets. The codes are available at https://github.com/TmacMai/Multimodal-Information-Bottleneck. Sijie Mai, Haifeng Hu 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Multimodal Graph for Unaligned Multimodal Sequence Analysis via Graph Convolution and Graph PoolingabstractMultimodal sequence analysis aims to draw inferences from visual, language, and acoustic sequences. A majority of existing works focus on the aligned fusion of three modalities to explore inter-modal interactions, which is impractical in real-world scenarios. To overcome this issue, we seek to focus on analyzing unaligned sequences, which is still relatively underexplored and also more challenging. We propose Multimodal Graph, whose novelty mainly lies in transforming the sequential learning problem into graph learning problem. The graph-based structure enables parallel computation in time dimension (as opposed to recurrent neural network) and can effectively learn longer intra- and inter-modal temporal dependency in unaligned sequences. First, we propose multiple ways to construct the adjacency matrix for sequence to perform sequence to graph transformation. To learn intra-modal dynamics, a graph convolution network is employed for each modality based on the defined adjacency matrix. To learn inter-modal dynamics, given that the unimodal sequences are unaligned, the commonly considered word-level fusion does not pertain. To this end, we innovatively devise graph pooling algorithms to automatically explore the associations between various time slices from different modalities and learn high-level graph representation hierarchically. Multimodal Graph outperforms state-of-the-art models on three datasets under the same experimental setting. Sijie Mai, Songlong Xing, Jiaxuan He, Haifeng Hu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2022 | Curriculum Learning Meets Weakly Supervised Multimodal Correlation LearningabstractIn the field of multimodal sentiment analysis (MSA), a few studies have leveraged the inherent modality correlation information stored in samples for self-supervised learning.However, they feed the training pairs in a random order without consideration of difficulty.Without human annotation, the generated training pairs of self-supervised learning often contain noise.If noisy or hard pairs are used for training at the easy stage, the model might be stuck in bad local optimum.In this paper, we inject curriculum learning into weakly supervised modality correlation learning.The weakly supervised correlation learning leverages the label information to generate scores for negative pairs to learn a more discriminative embedding space, where negative pairs are defined as two unimodal embeddings from different samples.To assist the correlation learning, we feed the training pairs to the model according to difficulty by the proposed curriculum learning, which consists of elaborately designed scoring and feeding functions.The scoring function computes the difficulty of pairs using pre-trained and current correlation predictors, where the pairs with large losses are defined as hard pairs.Notably, the hardest pairs are discarded in our algorithm, which are assumed as noisy pairs.Moreover, the feeding function takes the difference of correlation losses as feedback to determine the feeding actions ('stay', 'step back', or 'step forward').The proposed method reaches state-of-the-art performance on MSA. Sijie Mai, Haifeng Hu 0001 |
EMNLP | 1 |
| 2022 | Communicative Subgraph Representation Learning for Multi-Relational Inductive Drug-Gene Interaction PredictionabstractIlluminating the interconnections between drugs and genes is an important topic in drug development and precision medicine. Currently, computational predictions of drug-gene interactions mainly focus on the binding interactions without considering other relation types like agonist, antagonist, etc. In addition, existing methods either heavily rely on high-quality domain features or are intrinsically transductive, which limits the capacity of models to generalize to drugs/genes that lack external information or are unseen during the training process. To address these problems, we propose a novel Communicative Subgraph representation learning for Multi-relational Inductive drug-Gene interactions prediction (CoSMIG), where the predictions of drug-gene relations are made through subgraph patterns, and thus are naturally inductive for unseen drugs/genes without retraining or utilizing external domain features. Moreover, the model strengthened the relations on the drug-gene graph through a communicative message passing mechanism. To evaluate our method, we compiled two new benchmark datasets from DrugBank and DGIdb. The comprehensive experiments on the two datasets showed that our method outperformed state-of-the-art baselines in the transductive scenarios and achieved superior performance in the inductive ones. Further experimental analysis including LINCS experimental validation and literature verification also demonstrated the value of our model. Jiahua Rao, Shuangjia Zheng, Sijie Mai, Yuedong Yang |
IJCAI | 3 |
| 2022 | Contextual relation embedding and interpretable triplet capsule for inductive relation prediction
Sijie Mai, Haifeng Hu 0001 |
Neurocomputing | 2 |
| 2022 | Dynamic graph dropout for subgraph-based relation prediction
Sijie Mai, Shuangjia Zheng, Yuedong Yang, Haifeng Hu 0001 |
Knowl. Based Syst. | 1 |
| 2022 | Multi-Fusion Residual Memory Network for Multimodal Human Sentiment ComprehensionabstractMultimodal human sentiment comprehension refers to recognizing human affection from multiple modalities. There exist two key issues for this problem. First, it is difficult to explore time-dependent interactions between modalities and focus on the important time steps. Second, processing the long fused sequence of utterances is susceptible to the forgetting problem due to the long-term temporal dependency. In this article, we introduce a hierarchical learning architecture to classify utterance-level sentiment. To address the first issue, we perform time-step level fusion to generate fused features for each time step, which explicitly models time-restricted interactions by incorporating information across modalities at the same time step. Furthermore, based on the assumption that acoustic features directly reflect emotional intensity, we pioneer emotion intensity attention to focus on the time steps where emotion changes or intense affections take place. To handle the second issue, we propose Residual Memory Network (RMN) to process the fused sequence. RMN utilizes some techniques such as directly passing the previous state into the next time step, which helps to retain the information from many time steps ago. We show that our method achieves state-of-the-art performance on multiple datasets. Results also suggest that RMN yields competitive performance on sequence modeling tasks. Sijie Mai, Haifeng Hu 0001, Songlong Xing |
IEEE Trans. Affect. Comput. | 1 |
| 2022 | Adapted Dynamic Memory Network for Emotion Recognition in ConversationabstractIn this article, we address Emotion Recognition in Conversation (ERC) where conversational data are presented in a multimodal setting. Psychological evidence shows that self and inter-speaker influence are two central factors to emotion dynamics in conversation. State-of-the-art models do not effectively synthesise these two factors. Therefore, we propose an Adapted Dynamic Memory Network (A-DMN) where self and inter-speaker influences are modelled individually and further synthesised oriented towards the current utterance. Specifically, we model the dependency of the constituent utterances in a dialogue video using a global RNN to capture inter-speaker influence. Likewise, each speaker is assigned an RNN to capture their self influence. Afterwards, an Episodic Memory Module is devised to extract contexts for self and inter-speaker influence and synthesise them to update the memory. This process repeats itself for multiple passes until a refined representation is obtained and used for final prediction. Additionally, we explore cross-modal fusion in the context of multimodal ERC, and propose a convolution-based method which proves effective in extracting local interactions and computationally efficient. Extensive experiments demonstrate that A-DMN outperforms the state-of-the-art models on benchmark datasets. Songlong Xing, Sijie Mai, Haifeng Hu 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2022 | Interpretable Multimodal Capsule FusionabstractWith the development of social networking platform, multimodal sentiment analysis has become increasingly prominent. Existing models focus on capturing intramodal and intermodal interactions to produce effective modality representations. However, they overlook the study of interpretability which reveals how modalities interact with each other and which modality contributes most to the final prediction. In this paper, we propose an interpretable model called Interpretable Multimodal Capsule Fusion (IMCF) which integrates routing mechanism of Capsule Network (CapsNet) and Long Short-Term Memory (LSTM) to produce refined modality representations and provide interpretation. By constructing features of different modalities into input sequence, we are able to obtain highly expressive representation of intermodal dynamics due to the strong ability of LSTM to produce representation of sequence. As routing mechanism is applied during modality fusion and prediction stages, the value of routing coefficient can reveal the contributions of different modalities or dynamics, which provides interpretation. Meanwhile, routing mechanism can iteratively adjust the information flows of different modalities, which makes the process of modality fusion more reasonable. The experimental results show that our model achieves competitive performance on two benchmark datasets with effective modality fusion by LSTM and interpretation provided by routing mechanism. Sijie Mai, Haifeng Hu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | A Unimodal Representation Learning and Recurrent Decomposition Fusion Structure for Utterance-Level Multimodal Embedding LearningabstractLearning a unified embedding for utterance-level video attracts significant attention recently due to the rapid development of social media and its broad applications. An utterance normally contains not only spoken language but also the nonverbal behaviors such as facial expressions and vocal patterns. Instead of directly learning utterance embedding based on low-level features, we firstly explore high-level representation for each modality separately via an unimodal representation learning gyroscope structure. In this way, the learnt unimodal representations are more representative and contain more abstract semantic information. In the gyroscope structure, we introduce multi-scale kernel learning, ‘channel expansion’ and ‘channel fusion’ operations to explore high-level features both spatially and channelwise. Another insight of our method lies in that we fuse representations of all modalities to obtain a unified embedding by interpreting fusion procedure as the flow of inter-modality information between various modalities, which is more specialized in terms of the information to be fused and the fusion process. Specifically, considering that each modality carries modality-specific and cross-modality interactions, we innovate to decompose unimodal representations into intra- and inter-modality dynamics using gating mechanism, and further fuse the inter-modality dynamics by passing them from previous modalities to the following one using a recurrent neural fusion architecture. Extensive experiments demonstrate that our method achieves state-of-the-art performance on multiple benchmark datasets. Sijie Mai, Haifeng Hu 0001, Songlong Xing |
IEEE Trans. Multim. | 1 |
| 2021 | Communicative Message Passing for Inductive Relation ReasoningabstractRelation prediction for knowledge graphs aims at predicting missing relationships between entities. Despite the importance of inductive relation prediction, most previous works are limited to a transductive setting and cannot process previously unseen entities. The recent proposed subgraph-based relation reasoning models provided alternatives to predict links from the subgraph structure surrounding a candidate triplet inductively. However, we observe that these methods often neglect the directed nature of the extracted subgraph and weaken the role of relation information in the subgraph modeling. As a result, they fail to effectively handle the asymmetric/anti-symmetric triplets and produce insufficient embeddings for the target triplets. To this end, we introduce a Communicative Message Passing neural network for Inductive reLation rEasoning, CoMPILE, that reasons over local directed subgraph structures and has a vigorous inductive bias to process entity-independent semantic relations. In contrast to existing models, CoMPILE strengthens the message interactions between edges and entitles through a communicative kernel and enables a sufficient flow of relation information. Moreover, we demonstrate that CoMPILE can naturally handle asymmetric/anti-symmetric relations without the need for explosively increasing the number of model parameters by extracting the directed enclosing subgraphs. Extensive experiments show substantial performance gains in comparison to state-of-the-art methods on commonly used benchmark datasets with variant inductive settings. Sijie Mai, Shuangjia Zheng, Yuedong Yang, Haifeng Hu 0001 |
AAAI | 1 |
| 2021 | Graph Capsule Aggregation for Unaligned Multimodal SequencesabstractHumans express their opinions and emotions through multiple modalities which mainly consist of textual, acoustic and visual modalities. Prior works on multimodal sentiment analysis mostly apply Recurrent Neural Network (RNN) to model aligned multimodal sequences. However, it is unpractical to align multimodal sequences due to different sample rates for different modalities. Moreover, RNN is prone to the issues of gradient vanishing or exploding and it has limited capacity of learning long-range dependency which is the major obstacle to model unaligned multimodal sequences. In this paper, we introduce Graph Capsule Aggregation (GraphCAGE) to model unaligned multimodal sequences with graph-based neural model and Capsule Network. By converting sequence data into graph, the previously mentioned problems of RNN are avoided. In addition, the aggregation capability of Capsule Network and the graph-based structure enable our model to be interpretable and better solve the problem of long-range dependency. Experimental results suggest that GraphCAGE achieves state-of-the-art performance on two benchmark datasets with representations refined by Capsule Network and interpretation provided. Sijie Mai, Haifeng Hu 0001 |
ICMI | 2 |
| 2021 | A Unimodal Reinforced Transformer With Time Squeeze Fusion for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis refers to inferring sentiment from language, acoustic, and visual sequences. Previous studies focus on analyzing aligned sequences, while the unaligned sequential analysis is more practical in real-world scenarios. Due to the long-time dependency hidden in the multimodal unaligned sequence and time alignment information is not provided, exploring the time-dependent interactions within unaligned sequences is more challenging. To this end, we introduce the time squeeze fusion to automatically explore the time-dependent interactions by modeling the unimodal and multimodal sequences from the perspective of compressing the time dimension. Moreover, prior methods tend to fuse unimodal features into a multimodal embedding, based on which sentiment is inferred. However, we argue that the unimodal information may be lost or the generated multimodal embedding may be redundant. Addressing this issue, we propose a unimodal reinforced Transformer to progressively attend and distill unimodal information from the multimodal embedding, which enables the multimodal embedding to highlight the discriminative unimodal information. Extensive experiments suggest that our model reaches state-of-the-art performance in terms of accuracy and F1 score on MOSEI dataset. Jiaxuan He, Sijie Mai, Haifeng Hu 0001 |
IEEE Signal Process. Lett. | 2 |
| 2021 | Learning to Balance the Learning Rates Between Various Modalities via Adaptive Tracking FactorabstractMultimodal networks with richer information contents should always outperform the unimodal counterparts. In our experiment, however, we observe that this is not always the case. Prior efforts on multimodal tasks mainly tend to design a uniform optimization algorithm for all modalities, and yet only obtain a sub-optimal multimodal representation with the fusion of under-optimized unimodal representations, which are still challenged by performance drop on multimodal networks caused by heterogeneity among modalities. In this work, to remove the slowdowns in performance on multimodal tasks, we decouple the learning procedures of unimodal and multimodal networks by dynamically balancing the learning rates for various modalities, so that the modality-specific optimization algorithm for each modality can be obtained. Specifically, the adaptive tracking factor (ATF) is introduced to adjust the learning rate for each modality on a real-time basis. Furthermore, adaptive convergent equalization (ACE) and bilevel directional optimization (BDO) are proposed to equalize and update the ATF, avoiding sub-optimal unimodal representations due to overfitting or underfitting. Extensive experiments on multimodal sentiment analysis demonstrate that our method achieves superior performance. Sijie Mai, Haifeng Hu 0001 |
IEEE Signal Process. Lett. | 2 |
| 2021 | Analyzing Multimodal Sentiment Via Acoustic- and Visual-LSTM With Channel-Aware Temporal Convolution NetworkabstractThe emotion of human is always expressed in a multimodal perspective. Analyzing multimodal human sentiment remains challenging due to the difficulties of the interpretation in inter-modality dynamics. Mainstream multimodal learning architectures tend to design various fusion strategies to learn inter-modality interactions, which barely consider the fact that the language modality is far more important than the acoustic and visual modalities. In contrast, we learn inter-modality dynamics in a different perspective via acoustic- and visual-LSTMs where language features play dominant role. Specifically, inside each LSTM variant, a well-designed gating mechanism is introduced to enhance the language representation via the corresponding auxiliary modality. Furthermore, in the unimodal representation learning stage, instead of using RNNs, we introduce `channel-aware' temporal convolution network to extract high-level representations for each modality to explore both temporal and channel-wise interdependencies. Extensive experiments demonstrate that our approach achieves very competitive performance compared to the state-of-the-art methods on three widely-used benchmarks for multimodal sentiment analysis and emotion recognition. Sijie Mai, Songlong Xing, Haifeng Hu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2020 | Modality to Modality Translation: An Adversarial Representation Learning and Graph Fusion Network for Multimodal FusionabstractLearning joint embedding space for various modalities is of vital importance for multimodal fusion. Mainstream modality fusion approaches fail to achieve this goal, leaving a modality gap which heavily affects cross-modal fusion. In this paper, we propose a novel adversarial encoder-decoder-classifier framework to learn a modality-invariant embedding space. Since the distributions of various modalities vary in nature, to reduce the modality gap, we translate the distributions of source modalities into that of target modality via their respective encoders using adversarial training. Furthermore, we exert additional constraints on embedding space by introducing reconstruction loss and classification loss. Then we fuse the encoded representations using hierarchical graph neural network which explicitly explores unimodal, bimodal and trimodal interactions in multi-stage. Our method achieves state-of-the-art performance on multiple datasets. Visualization of the learned embeddings suggests that the joint embedding space learned by our method is discriminative. Sijie Mai, Haifeng Hu 0001, Songlong Xing |
AAAI | 1 |
| 2020 | Locally Confined Modality Fusion Network With a Global Perspective for Multimodal Human Affective ComputingabstractIn this paper, we propose a novel multimodal fusion framework, called the locally confined modality fusion network (LMFN), that contains a bidirectional multiconnected LSTM (BM-LSTM) to address the multimodal human affective computing problem. In the LMFN, we introduce a generic fusion structure that explores both local and global fusion to obtain an integral comprehension of information. Specifically, we partition the feature vector corresponding to each modality into multiple segments and learn every local interaction through a tensor fusion procedure. Global interaction is then modeled by learning the dependence between local tensors via an originally designed BM-LSTM architecture, establishing a direct connection of cells and states of local tensors that are several time steps apart. With the LMFN, we achieve advantages over other methods in the following aspects: 1) local interactions are successfully modeled using a feasible vector segmentation procedure that can explore cross-modal dynamics in a more specialized manner; 2) global interactions are modeled to obtain an integral view of multimodal information using BM-LSTM, which guarantees an adequate flow of information; and 3) our general fusion structure is highly extendable by applying other local and global fusion methods. Experiments show that the LMFN yields state-of-the-art results. Moreover, the LMFN achieves higher efficiency compared to other models by applying the outer product as the fusion method. Sijie Mai, Songlong Xing, Haifeng Hu 0001 |
IEEE Trans. Multim. | 1 |
| 2019 | Divide, Conquer and Combine: Hierarchical Feature Fusion Network with Local and Global Perspectives for Multimodal Affective ComputingabstractWe propose a general strategy named 'divide, conquer and combine' for multimodal fusion.Instead of directly fusing features at holistic level, we conduct fusion hierarchically so that both local and global interactions are considered for a comprehensive interpretation of multimodal embeddings.In the 'divide' and 'conquer' stages, we conduct local fusion by exploring the interaction of a portion of the aligned feature vectors across various modalities lying within a sliding window, which ensures that each part of multimodal embeddings are explored sufficiently.On its basis, global fusion is conducted in the 'combine' stage to explore the interconnection across local interactions, via an Attentive Bi-directional Skipconnected LSTM that directly connects distant local interactions and integrates two levels of attention mechanism.In this way, local interactions can exchange information sufficiently and thus obtain an overall view of multimodal information.Our method achieves state-ofthe-art performance on multimodal affective computing with higher efficiency. Sijie Mai, Haifeng Hu 0001, Songlong Xing |
ACL (1) | 1 |
| 2019 | Attentive matching network for few-shot learning
Sijie Mai, Haifeng Hu 0001 |
Comput. Vis. Image Underst. | 1 |