Guimin Hu

dblp:275/9112 · DBLP profile ↗
← Back
16ranked-venue papers
8as first author
15since 2021 · last 2026
0000-0001-8364-3076ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 7 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 S³-MSD: Large Vision-Language Model for Explainable and Generalizable Multi-modal Sarcasm Detection
abstract
Multimodal sarcasm detection (MSD) aims to identify sarcasm polarity from diverse modalities (i.e., image–text pairs), a task that has received increasing attention. While significant progress has been made, existing approaches still face two major issues: lack of explainability and weak generalizability. In this paper, we introduce a new large vision–language model (LVLM) dubbed S³-MSD for explainable and generalizable MSD through three key components. For explainability, we develop (1) a self-training paradigm that automatically bootstraps answers with explanations, and (2) a self-calibrating mechanism that rectifies flawed explanations. For generalizability, we design (3) a self-focusing module that amplifies visual semantic entities through preference optimization, thereby mitigating textual over-reliance. Experimental results on both in-distribution and out-of-distribution (OOD) benchmarks demonstrate that S³-MSD consistently outperforms state-of-the-art methods in detection performance. Furthermore, the proposed S³-MSD provides persuasive explanations, as verified by both quantitative metrics and human evaluations.
Zhihong Zhu 0001, Fan Zhang 0111, Yunyan Zhang, Jinghan Sun, Guimin Hu, Hao Wu 0094, Yuyan Chen, Xian Wu 0001
AAAI5
2026 CMCTS: A Constrained Monte Carlo Tree Search framework for mathematical reasoning in large language model
Qingwen Lin, Guimin Hu, Zijian Li 0001, Zhifeng Hao 0004, Keli Zhang, Ruichu Cai
Appl. Intell.3
2025 Explicitly Guided Difficulty-Controllable Visual Question Generation
abstract
Visual question generation (VQG) aims to generate questions from images automatically. While existing studies primarily focus on the quality of generated questions, such as fluency and relevance, the difficulty of the questions is also a crucial factor in assessing their quality. Question difficulty directly impacts the effectiveness of VQG systems in applications like education and human-computer interaction, where appropriately challenging questions can stimulate learning interest and improve interaction experiences. However, accurately defining and controlling question difficulty is a challenging task due to its multidimensional and subjective nature. In this paper, we propose a new definition of the difficulty of questions, i.e., being positively correlated with the number of reasoning steps required to answer a question. For our definition, we construct a corresponding dataset and propose a benchmark as a foundation for future research. Our benchmark is designed to progressively increase the reasoning steps involved in generating questions. Specifically, we first extract the relationships among objects in the image to form a reasoning chain, then gradually increase the difficulty by rewriting the generated question to include more reasoning sub-chains. Experimental results on our constructed dataset show that our benchmark significantly outperforms existing baselines in controlling the reasoning chains of generated questions, producing questions with varying difficulty levels.
Jiayuan Xie, Mengqiu Cheng, Xinting Zhang, Yi Cai 0001, Guimin Hu, Mengying Xie, Qing Li 0001
AAAI5
2025 Debiasing Multilingual LLMs in Cross-lingual Latent Space
abstract
Debiasing techniques such as SentDebias aim to reduce bias in large language models (LLMs).Previous studies have evaluated their cross-lingual transferability by directly applying these methods to LLM representations, revealing their limited effectiveness across languages.In this work, we therefore propose to perform debiasing in a joint latent space rather than directly on LLM representations.We construct a well-aligned cross-lingual latent space using an autoencoder trained on parallel TED talk scripts.Our experiments with Aya-expanse and two debiasing techniques across four languages (English, French, German, Dutch) demonstrate that a) autoencoders effectively construct a well-aligned cross-lingual latent space, and b) applying debiasing techniques in the learned cross-lingual latent space significantly improves both the overall debiasing performance and cross-lingual transferability.
Qiwei Peng 0003, Guimin Hu, Yekun Chai, Anders Søgaard
EMNLP2
2025 PgM: Partitioner Guided Modal Learning Framework
abstract
Multimodal learning benefits from multiple modal information, and each learned modal representations can be divided into uni-modal that can be learned from uni-modal training and paired-modal features that can be learned from cross-modal interaction. Building on this perspective, we propose a partitioner-guided modal learning framework, PgM, which consists of the modal partitioner, uni-modal learner, paired-modal learner, and uni-paired modal decoder. Modal partitioner segments the learned modal representation into uni-modal and paired-modal features. Modal learner incorporates two dedicated components for uni-modal and paired-modal learning. Uni-paired modal decoder reconstructs modal representation based on uni-modal and paired-modal features. PgM offers three key benefits: 1) thorough learning of uni-modal and paired-modal features, 2) flexible distribution adjustment for uni-modal and paired-modal representations to suit diverse downstream tasks, and 3) different learning rates across modalities and partitions. Extensive experiments demonstrate the effectiveness of PgM across four multimodal tasks and further highlight its transferability to existing models. Additionally, we visualize the distribution of uni-modal and paired-modal features across modalities and tasks, offering insights into their respective contributions.
Guimin Hu, Yi Xin 0003, Lijie Hu, Zhihong Zhu 0001, Hasti Seifi
ACM Multimedia1
2024 Towards Multi-modal Sarcasm Detection via Disentangled Multi-grained Multi-modal Distilling
abstract
Multi-modal sarcasm detection aims to identify whether a given sample with multi-modal information (i.e., text and image) is sarcastic, which has received increasing attention due to the rapid growth of multi-modal posts on modern social media. However, mainstream models process the input of each modality in a holistic manner, resulting in redundant and unrefined information. Moreover, the representations of different modalities are entangled in one common latent space to perform complex cross-modal interactions, neglecting the heterogeneity and distribution gap of different modalities. To address these issues, we propose a novel framework DMMD (short for Disentangled Multi-grained Multi-modal Distilling) for multi-modal sarcasm detection, which conducts multi-grained knowledge distilling (i.e., intra-subspace and inter-subspace) based on the disentangled multi-modal representations. Concretely, the representations of each modality are disentangled explicitly into modality-agnostic/specific subspaces. Then we transfer cross-modal knowledge by conducting intra-subspace knowledge distilling in a self-adaptive pattern. We also apply mutual learning to regularize the underlying inter-subspace consistency. Extensive experiments on a commonly used benchmark demonstrate the efficacy of our DMMD over cutting-edge methods. More encouragingly, visualization results indicate the multi-modal representations display meaningful distributional patterns, and we hope it will be helpful for the community of multi-modal knowledge transfer.
Zhihong Zhu 0001, Xuxin Cheng, Guimin Hu, Yaowei Li 0001, Zhiqi Huang 0001, Yuexian Zou
LREC/COLING3
2024 FoodieQA: A Multimodal Dataset for Fine-Grained Understanding of Chinese Food Culture
abstract
Wenyan Li, Crystina Zhang, Jiaang Li, Qiwei Peng, Raphael Tang, Li Zhou, Weijia Zhang, Guimin Hu, Yifei Yuan, Anders Søgaard, Daniel Hershcovich, Desmond Elliott. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Wenyan Li 0001, Xinyu Zhang 0018, Jiaang Li 0002, Qiwei Peng 0003, Raphael Tang, Li Zhou 0010, Weijia Zhang 0004, Guimin Hu, Yifei Yuan 0002, Anders Søgaard, Daniel Hershcovich, Desmond Elliott
EMNLP8
2024 TFCD: Towards Multi-modal Sarcasm Detection via Training-Free Counterfactual Debiasing
Zhihong Zhu 0001, Xianwei Zhuang, Yunyan Zhang, Derong Xu, Guimin Hu, Xian Wu 0001, Yefeng Zheng 0001
IJCAI5
2024 V-PETL Bench: A Unified Visual Parameter-Efficient Transfer Learning Benchmark
abstract
Parameter-efficient transfer learning (PETL) methods show promise in adapting a pre-trained model to various downstream tasks while training only a few parameters. In the computer vision (CV) domain, numerous PETL algorithms have been proposed, but their direct employment or comparison remains inconvenient. To address this challenge, we construct a Unified Visual PETL Benchmark (V-PETL Bench) for the CV domain by selecting 30 diverse, challenging, and comprehensive datasets from image recognition, video action recognition, and dense prediction tasks. On these datasets, we systematically evaluate 25 dominant PETL algorithms and open-source a modular and extensible codebase for fair evaluation of these algorithms. V-PETL Bench runs on NVIDIA A800 GPUs and requires approximately 310 GPU days. We release all the benchmark, making it more efficient and friendly to researchers. Additionally, V-PETL Bench will be continuously updated for new PETL algorithms and CV tasks.
Yi Xin 0003, Xuyang Liu 0002, Yuntao Du 0001, Haodi Zhou, Christina E. Lee, Junlong Du, Haozhe Wang 0002, Mingcai Chen, Ting Liu 0018, Guimin Hu, Zhongwei Wan, Rongchao Zhang, Aoxue Li, Mingyang Yi, Xiaohong Liu 0001
NeurIPS12
2024 Unifying emotion-oriented and cause-oriented predictions for emotion-cause pair extraction
Guimin Hu, Yi Zhao 0007, Guangming Lu 0002
Neural Networks1
2024 Improving Representation With Hierarchical Contrastive Learning for Emotion-Cause Pair Extraction
abstract
Emotion-cause pair extraction (ECPE) aims to extract emotions and their corresponding cause from a document. The previous works have made great progress. However, there exist two major issues in existing works. First, most existing works mainly focus on the semantic relation between the emotion clause and cause clause, ignoring their inner statistical relation in representation space. Second, the existing works are sensitive to the relative position between the emotion clause and cause clause, which damages the model's robustness. To address the two issues, we propose a hierarchical contrastive learning framework (HCL-ECPE), which hierarchically performs contrastive learning on representation from two levels. The first level is inter-clause contrastive learning (ICCL), which performs between emotion clause and cause clause through mutual information maximization. The second level is intra-pair contrastive learning (IPCL), which performs between clause representation and pair representation through contrastive predictive coding (CPC). HCL-ECPE integrates ICCL and IPCL modules to explore the statistical relations between the emotion clause, cause clause, and their constructed emotion-cause pair from the perspective of mutual information, thereby improving the model performance and robustness. Experimental results on two public datasets, ECPED and RECCON, demonstrate that HCL-ECPE outperforms the most competitive baselines. Furthermore, ICCL and IPCL are orthogonal to the existing model, and introducing them into the current models updates state-of-the-art performance.
Guimin Hu, Yi Zhao 0007, Guangming Lu 0002
IEEE Trans. Affect. Comput.1
2023 Emotion Prediction Oriented Method With Multiple Supervisions for Emotion-Cause Pair Extraction
abstract
Emotion-cause pair extraction (ECPE) task aims to extract all the pairs of emotions and their causes from an unannotated emotion text. The previous works usually extract the emotion-cause pairs from two perspectives of emotion and cause. However, emotion extraction is more crucial to the ECPE task than cause extraction. Motivated by this analysis, we propose an end-to-end emotion-cause extraction approach oriented toward emotion prediction (EPO-ECPE), aiming to fully exploit the potential of emotion prediction to enhance emotion-cause pair extraction. Considering the strong dependence between emotion prediction and emotion-cause pair extraction, we propose a synchronization mechanism to share their improvement in the training process. That is, the improvement of emotion prediction can facilitate the emotion-cause pair extraction, and then the results of emotion-cause pair extraction can also be used to improve the accuracy of emotion prediction simultaneously. For the emotion-cause pair extraction, we divide it into genuine pair supervision and fake pair supervision, where the genuine pair supervision learns from the pairs with more possibility to be emotion-cause pairs. In contrast, fake pair supervision learns from other pairs. In this way, the emotion-cause pairs can be extracted directly from the genuine pair, thereby reducing the difficulty of extraction. Experimental results show that our approach outperforms the 13 compared systems and achieves new state-of-the-art performance.
Guimin Hu, Yi Zhao 0007, Guangming Lu 0002
IEEE ACM Trans. Audio Speech Lang. Process.1
2022 UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition
abstract
Multimodal sentiment analysis (MSA) and emotion recognition in conversation (ERC) are key research topics for computers to understand human behaviors.From a psychological perspective, emotions are the expression of affect or feelings during a short period, while sentiments are formed and held for a longer period.However, most existing works study sentiment and emotion separately and do not fully exploit the complementary knowledge behind the two.In this paper, we propose a multimodal sentiment knowledge-sharing framework (UniMSE) that unifies MSA and ERC tasks from features, labels, and models.We perform modality fusion at the syntactic and semantic levels and introduce contrastive learning between modalities and samples to better capture the difference and consistency between sentiments and emotions.Experiments on four public benchmark datasets, MOSI, MOSEI, MELD, and IEMO-CAP, demonstrate the effectiveness of the proposed method and achieve consistent improvements compared with state-of-the-art methods.
Guimin Hu, Ting-En Lin, Yi Zhao 0007, Guangming Lu 0002, Yuchuan Wu
EMNLP1
2022 An exploration of mutual information based on emotion-cause pair extraction
Guimin Hu, Yi Zhao 0007, Guangming Lu 0002, Fanghao Yin, Jiashan Chen
Knowl. Based Syst.1
2021 FSS-GCN: A graph convolutional networks with fusion of semantic and structure for emotion cause analysis
Guimin Hu, Guangming Lu 0002, Yi Zhao 0007
Knowl. Based Syst.1
2020 Emotion-Cause Joint Detection: A Unified Network with Dual Interaction for Emotion Cause Analysis
Guimin Hu, Guangming Lu 0002, Yi Zhao 0007
NLPCC (1)1