EDBT 2026 Demo / reviewers in the wild / expert
Minjie Ren
dblp:253/8946 · also Min-Jie Ren
· DBLP profile ↗
10ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0002-6080-8074ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FACT: A Simple and Efficient Framework for Active FinetuningabstractThe main goal of active finetuning is to improve a pretrained model's performance on a specific task or domain by finetuning it with carefully selected informative or challenging data. Previous research has predominantly focused on the active aspect (i.e., data selection) while uniformly employing full finetuning for model adaptation, which inevitably distorts pretrained features due to distribution shift. This issue becomes particularly pronounced when the model size is large relative to the finetuning data quantity, leading to heightened overfitting risks. To address this critical gap, we formally outline the FiAF task that emphasizes systematic exploration of finetuning methodologies in active learning. We propose FACT, a three-phase hierarchical finetuning framework featuring both efficiency and simplicity, specifically designed for active finetuning scenarios. Our comprehensive experiments span: 1) Three major dataset categories encompassing classic (CIFAR10, CIFAR100, ImageNet-1k), imbalanced (CIFAR10-LT, CIFAR100-LT), and fine-grained (StanfordCars, FGVCAircraft) image classification datasets, each evaluated under 3-5 distinct sampling ratios; 2) Diverse pretrained architectures including Convolutional Neural Network (ConvNeXt), Vision Transformer (ViT), and Vision LSTM (ViL) networks; 3) A systematic investigation of frozen feature augmentation (FroFA) strategies. 4) A comprehensive and rigorous analysis of efficiency and generalizability. The results demonstrate significant improvements with strong generalization and robustness. Notably, under low sampling ratios, our framework achieves remarkable performance gains of over 20% on the ViT model for CIFAR10, CIFAR100, and ImageNet-1k benchmarks. This systematic approach establishes new state-of-the-art performance while maintaining parameter efficiency, proving particularly effective when labeled data is scarce. Wenshuai Xu, You Song, Yuzhuo Cui, Minjie Ren, Qingjie Liu 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | KA-MIN: Knowledge-Aware Multimodal Interaction Network for Emotion Recognition in ConversationabstractEmotion recognition in conversations (ERC) has garnered significant attention for its critical role in human-computer interaction systems. ERC benefits from multimodal data, which offers diverse perspectives on emotional states, and commonsense knowledge (CSK), which enriches the context by incorporating real-world understanding of human behavior. However, existing ERC studies have not fully exploited the potential of multimodal-CSK interactions for complementary information learning from these sources. To address this, we innovatively propose a Knowledge-Aware Multimodal Interaction Network (KA-MIN). KA-MIN is designed to capture complementary emotional information from CSK-multimodal interactions, thereby facilitating the ERC task. To achieve this, KA-MIN begins by combining six relation types of CSK, leveraging their differences between multimodal emotional information. The fused CSK features are then refined to incorporate context and emotional information using multimodal contextual guidance. Subsequently, we construct a novel knowledge-aware multimodal graph structure that allows the CSK information to interact with multimodal information, leading to more comprehensive multimodal and context modeling. During the graph learning process, the CSK-multimodal interactions capture the complementary emotional information between CSK and multimodal features. Finally, we dynamically fuse the multimodal emotional information with the informative CSK and textual guidance to obtain the final utterance representations, which encompass effective emotional information from both multimodal and CSK features. Extensive experiments on two popular multimodal ERC datasets demonstrate the superiority and effectiveness of the proposed KA-MIN framework. Minjie Ren, Xiangdong Huang 0002, Jing Liu 0002, Zan Gao 0001, Yuting Su 0001, Anan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Multi-loop graph convolutional network for multimodal conversational emotion recognition
Minjie Ren, Xiangdong Huang 0002, Wenhui Li 0001, Jing Liu 0002 |
J. Vis. Commun. Image Represent. | 1 |
| 2023 | MALN: Multimodal Adversarial Learning Network for Conversational Emotion RecognitionabstractMultimodal emotion recognition in conversations (ERC) aims to identify the emotional state of constituent utterances expressed by multiple speakers in dialogue from multimodal data. Existing multimodal ERC approaches focus on modeling the global context of the dialogue and neglect to mine the characteristic information from the corresponding utterances expressed by the same speaker. Additionally, information from different modalities exhibits commonality and diversity for emotional expression. The commonality and diversity of multimodal information are compensated for each other but not effectively exploited in previous multimodal ERC works. To tackle these issues, we propose a novel Multimodal Adversarial Learning Network (MALN). MALN first mines the speaker’s characteristics from context sequences and then incorporate them with the unimodal features. Afterward, we design a novel adversarial module AMDM to exploit both commonality and diversity from the unimodal features. Finally, AMDM fuses different modalities to generate refined utterance representations for emotion classification. Extensive experiments are conducted on two public multimodal ERC datasets, IEMOCAP and MELD. Through the experiments, MALN shows its superiority over the state-of-the-art methods. Minjie Ren, Xiangdong Huang 0002, Jing Liu 0002, Ming Liu 0004, Xuanya Li, Anan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | I-GCN: Incremental Graph Convolution Network for Conversation Emotion DetectionabstractSentiment analysis and emotion detection in conversation are becoming hot topics in regard to several applications. With the development of the social robot, social network, and intelligent voice assistant, emotion detection is attracting more attention as a key component in these research fields. Many approaches have been proposed to handle this problem in recent years. However, these previous approaches focus on either the temporal change information of the conversation or the semantic correlation information of the dialogue but ignore the combination of temporal information and semantic correlation information. In this paper, we propose an incremental graph convolution network (I-GCN) to handle emotion detection in conversation. We first utilize the graph structure to represent conversation at different times, which can represent the semantic correlation information of utterances. Then, we apply the incremental graph structure to imitate the process of dynamic conversation, which can preserve the temporal change information of conversation. Especially, for the first step of the process, we creatively propose utterance-level GCN (U-GCN) and speaker-level GCN (S-GCN) to learn the features of utterances for emotion detection. U-GCN focuses on the correlations among utterances and applies the multi-head attention model to find latent correlation information among utterances, which aims to further enhance the guidance of semantic relevance for feature learning. S-GCN focuses on the correlation between speaker and utterances, which can provide a different angle to guide the feature learning of utterances. In the learning of model parameters, we constantly utilize the new utterances to fine-tune the parameters of GNN for enhancement of the contribution of temporal change information. Detailed evaluations of the proposed method on three published conversation corpuses demonstrate the great effectiveness of our approach over several conventional competitive baselines. Weizhi Nie, Rihao Chang, Minjie Ren, Yuting Su 0001, Anan Liu |
IEEE Trans. Multim. | 3 |
| 2022 | LR-GCN: Latent Relation-Aware Graph Convolutional Network for Conversational Emotion RecognitionabstractAs an intersection of artificial intelligence and human communication analysis, Emotion Recognition in Conversation (ERC) has attracted much research attention in recent years. Existing studies, however, are limited in adequately exploiting latent relations among the constituent utterances. In this paper, we address this issue by proposing a novel approach named Latent Relation-Aware Graph Convolutional Network (LR-GCN), where both speaker dependency of the interlocutors is leveraged and latent correlations among the utterances are captured for ERC. Specifically, we first establish a graph model to incorporate the context information and speaker dependency of the conversation. Afterward, the multi-head attention mechanism is introduced to explore the latent correlations among the utterances and generate a set of all-linked graphs. Here, aiming to simultaneously exploit the original modeled speaker dependency and the explored correlation information, we introduce a dense connection layer to capture more structural information of the generated graphs. Through a multi-branch graph network, we achieve a unified representation of each utterance for final prediction. Detailed evaluations on two benchmark datasets demonstrate LR-GCN outperforms the state-of-the-art approaches. Minjie Ren, Xiangdong Huang 0002, Wenhui Li 0001, Dan Song 0006, Weizhi Nie |
IEEE Trans. Multim. | 1 |
| 2021 | Interactive Multimodal Attention Network for Emotion Recognition in ConversationabstractIn this letter, we propose a novel Interactive Multimodal Attention Network (IMAN) for emotion recognition in conversations. IMAN introduces a cross-modal attention fusion module to capture cross-modal interactions of multimodal information, and employs a conversational modeling module to explore the context information and speaker dependency of the whole conversation. Concretely, the cross-modal attention fusion module captures the cross-modal interactions and complementary information among the pre-extracted unimodal features from textual, visual, acoustic modalities based on the cross-modal attention block. Afterward, the updated features from each modality are fused to concentrate more on the informative modality and achieve a refined feature for each constituent utterance. The conversational modeling module defines three different gated recurrent units (GRUs) with respect to the context information, the speaker dependency, and the emotional state of utterances. In this way, we exploit the speaker dependency and contextual information to obtain the emotional state of utterances for emotion classification. Empirical evaluations on the multimodal benchmark IEMOCAP dataset demonstrate that our IMAN achieves competitive performance compared to the state-of-the-art approaches. Minjie Ren, Xiangdong Huang 0002, Xiaoqi Shi, Weizhi Nie |
IEEE Signal Process. Lett. | 1 |
| 2021 | M-GCN: Multi-Branch Graph Convolution Network for 2D Image-based on 3D Model Retrievalabstract2D image based 3D model retrieval is a challenging research topic in the field of 3D model retrieval. The huge gap between two modalities - 2D image and 3D model, extremely constrains the retrieval performance. In order to handle this problem, we propose a novel multi-branch graph convolution network (M-GCN) to address the 2D image based 3D model retrieval problem. First, we compute the similarity between 2D image and 3D model based on visual information to construct one cross-modalities graph model, which can provide the original relationship between image and 3D model. However, this relationship is not accurate because of the difference of modalities. Thus, the multi-head attention mechanism is employed to generate a set of fully connected edge-weighted graphs, which can predict the hidden relationship between 2D image and 3D model to further strengthen the correlation for the embedding generation of nodes. Finally, we apply the max-pooling operation to fuse the multi-graphs information and generate the fusion embeddings of nodes for retrieval. To validate the performance of our method, we evaluated M-GCN on the MI3DOR dataset, Shrec 2018 track and Shrec 2014 track. The experimental results demonstrate the superiority of our proposed method over the state-of-the-art methods. Weizhi Nie, Minjie Ren, Anan Liu, Zhendong Mao 0001, Jie Nie |
IEEE Trans. Multim. | 2 |
| 2021 | C-GCN: Correlation Based Graph Convolutional Network for Audio-Video Emotion RecognitionabstractWith the development of both hardware and deep neural network technologies, tremendous improvements have been achieved in the performance of automatic emotion recognition (AER) based on the video data. However, AER is still a challenging task due to subtle expression, abstract concept of emotion and the representation of multi-modal information. Most proposed approaches focus on the multi-modal feature learning and fusion strategy, which pay more attention to the characteristic of a single video and ignore the correlation among the videos. To explore this correlation, in this paper, we propose a novel correlation-based graph convolutional network (C-GCN) for AER, which can comprehensively consider the correlation of the intra-class and inter-class videos for feature learning and information fusion. More specifically, we introduce the graph model to represent the correlation among the videos. This correlated information can help to improve the discrimination of node features in the progress of graph convolutional network. Meanwhile, the multi-head attention mechanism is applied to predict the hidden relationship among the videos, which can strengthen the inter-class correlation to improve the performance of classifiers. The C-GCN is evaluated on the AFEW datasets and eNTERFACE 05 dataset. The final experimental results demonstrate the superiority of our proposed method over the state-of-the-art methods. Weizhi Nie, Minjie Ren, Jie Nie, Sicheng Zhao |
IEEE Trans. Multim. | 2 |
| 2019 | Multi-modal Correlated Network for emotion recognition in speechabstractWith the growing demand of automatic emotion recognition system, emotion recognition is becoming more and more crucial for human–computer interaction (HCI) research. Recently, there is a continuous improvement in the performance of automatic emotion recognition due to the development of both hardware and deep learning methods. However, because of the abstract concept and multiple expressions of emotion, automatic emotion recognition is still a challenging task. In this paper, we propose a novel Multi-modal Correlated Network for emotion recognition aiming at exploiting the information from both audio and visual channels to achieve more robust and accurate detection. In the proposed method, the audio signals and visual signals are first preprocessed for the feature extraction. After preprocessing, we obtain the Mel-spectrograms, which can be treated as images, and the representative frames from visual segments. Then the Mel-spectrograms are fed to the convolutional neural network (CNN) to get the audio features and the representative frames are fed to the CNN and LSTM to get features. Specially, we employ the triplet loss to increase the differentiation of inter-class. Meanwhile, we propose a novel correlated loss to reduce the differentiation of intra-class. Finally, we apply the feature fusion method to fuse the audio and visual feature for emotion recognition classification. The experimental result on AEFW dataset demonstrates the correlation information of multiple modals is crucial for automatic emotion recognition and the proposed method can achieve the state-of-the-art performance on the classification task . Minjie Ren, Weizhi Nie, Anan Liu, Yuting Su 0001 |
Vis. Informatics | 1 |