VLDB 2026 Research / reviewers in the wild / expert
Shuzhen Li
dblp:172/9956
· DBLP profile ↗
20ranked-venue papers
5as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 2 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A coarse-to-fine dynamic layer pruning framework for parameter-efficient fine-tuning
Xin Zhang 0100, Shuzhen Li, Zhulin Liu, C. L. Philip Chen |
Neurocomputing | 2 |
| 2025 | A Parameter-Efficient and Fine-Grained Prompt Learning for Vision-Language ModelsabstractCurrent vision-language models (VLMs) understand complex vision-text tasks by extracting overall semantic information from largescale cross-modal associations.However, extracting from large-scale cross-modal associations often smooths out semantic details and requires large computations, limiting multimodal fine-grained understanding performance and efficiency.To address this issue, this paper proposes a detail-oriented prompt learning (DoPL) method for vision-language models to implement fine-grained multi-modal semantic alignment with merely 0.25M trainable parameters.According to the low-entropy information concentration theory, DoPL explores shared interest tokens from text-vision correlations and transforms them into alignment weights to enhance text prompt and vision prompt via detail-oriented prompt generation.It effectively guides the current frozen layer to extract fine-grained text-vision alignment cues.Furthermore, DoPL constructs detail-oriented prompt generation for each frozen layer to implement layer-by-layer localization of finegrained semantic alignment, achieving precise understanding in complex vision-text tasks.DoPL performs well in parameter-efficient finegrained semantic alignment with only 0.12% tunable parameters for vision-language models.The state-of-the-art results over the previous parameter-efficient fine-tuning methods and full fine-tuning approaches on six benchmarks demonstrate the effectiveness and efficiency of DoPL in complex multi-modal tasks. Yongbin Guo, Shuzhen Li, Zhulin Liu, Tong Zhang 0015, C. L. Philip Chen |
ACL (1) | 2 |
| 2025 | Incongruity-aware Tension Field Network for Multi-modal Sarcasm DetectionabstractMulti-modal sarcasm detection (MSD) identifies sarcasm and accurately understands users' real attitudes from text-image pairs.Most MSD researches explore the incongruity of textimage pairs as sarcasm information through consistency preference methods.However, these methods prioritize consistency over incongruity and blur incongruity information under their global feature aggregation mechanisms, leading to incongruity distortions and model misinterpretations.To address the above issues, this paper proposes a pioneering inconsistency preference method called incongruityaware tension field network (ITFNet) for multimodal sarcasm detection tasks.Specifically, ITFNet extracts effective text-image feature pairs in fact and sentiment perspectives.It then constructs a fact/sentiment tension field with discrepancy metrics to capture the contextual tone and polarized incongruity after the iterative learning of tension intensity, effectively highlighting incongruity information during such inconsistency preference learning.It further standardizes the polarized incongruity with reference to contextual tone to obtain standardized incongruity, effectively implementing instance standardization for unbiased decision-making in MSD.ITFNet performs well in extracting salient and standardized incongruity through an incongruity-aware tension field, significantly tackling incongruity distortions and cross-instance variance.Moreover, ITFNet achieves state-of-the-art performance surpassing LLaVA1.5-7B with only 17.3M trainable parameters, demonstrating its optimal performance-efficiency in multi-modal sarcasm detection tasks. Jiecheng Zhang, C. L. Philip Chen, Shuzhen Li, Tong Zhang 0015 |
ACL (1) | 3 |
| 2025 | An Orthogonal High-Rank Adaptation for Large Language ModelsabstractLow-rank adaptation (LoRA) efficiently adapts LLMs to downstream tasks by decomposing LLMs' weight update into trainable low-rank matrices for fine-tuning.However, the random low-rank matrices may introduce massive taskirrelevant information, while their recomposed form suffers from limited representation spaces under low-rank operations.Such dense and choked adaptation in LoRA impairs the adaptation performance of LLMs on downstream tasks.To address these challenges, this paper proposes OHoRA, an orthogonal high-rank adaptation for parameter-efficient fine-tuning on LLMs.According to the information bottleneck theory, OHoRA decomposes LLMs' pre-trained weight matrices into orthogonal basis vectors via QR decomposition and splits them into two low-redundancy high-rank components to suppress task-irrelevant information.It then performs dynamic rank-elevated recomposition through Kronecker product to generate expansive task-tailored representation spaces, enabling precise LLM adaptation and enhanced generalization.OHoRA effectively operationalizes the information bottleneck theory to decompose LLMs' weight matrices into low-redundancy high-rank components and recompose them in rank-elevated manner for more task-tailored representation spaces and precise LLM adaptation.Empirical evaluation shows OHoRA's effectiveness by outperforming LoRA and its variants and achieving comparable performance to full fine-tuning with only 0.0371% trainable parameters. Xin Zhang 0100, Guang-Ze Chen, Shuzhen Li, Zhulin Liu, C. L. Philip Chen, Tong Zhang 0015 |
EMNLP | 3 |
| 2025 | Label Feature Co-Learning for Facial and EEG Emotion RecognitionabstractRecognizing human emotions through behavioral and physiological signals is fundamental to overall health. However, since emotion occurs transiently, a semantics mismatch exists between the uniformly annotated label and multimodal temporal signals, leading to emotional ambiguity. Previous studies used the annotated label to guide feature learning, which makes it intractable to accurately identify emotional elicitation moments within each signal. Moreover, the inconsistency of specific elicitation moments across different signals complicates emotion recognition. The model hardly learns discriminative features due to emotional ambiguity, which weakens its ability to differentiate between emotions. To tackle the above challenges, this paper proposes a novel label feature co-learning model (LFCL) for emotion recognition through multimodal signals. Specifically, the LFCL leverages unimodal and multimodal information and adaptively generates instance-level emotion labels, thus precisely locating emotion elicitation moments within each signal. To promote emotion consistency across different signals, the LFCL incorporates a dynamic label calibration mechanism to balance the label generation process with historical information. Furthermore, to enhance the deep interaction between signals, the LFCL conducts multimodal interactive fusion to integrate multi-level multimodal features and extract global emotional information. The LFCL performs precise label-to-feature alignment to capture discriminative features of each signal, effectively alleviating emotional ambiguity and improving the ability to distinguish different emotions. Extensive experiments on three publicly available datasets demonstrate the effectiveness and generalization of the proposed model. Mingchen Cai, C. L. Philip Chen, Shuzhen Li, Tong Zhang 0015 |
IEEE Trans. Affect. Comput. | 3 |
| 2025 | From Disagreement to Unity: A Cascade Perspective Network for Subjective TasksabstractSubjective tasks involve annotating instances according to personal opinions, emotions, and feelings. The neural network model must learn about human thought and expression complexity. Handling instances with different annotation opinions is the main challenge in subjective tasks. Existing methods focus on majority voting or integrating opinions from a few assigned annotators, leading to biased decisions and limited performance. To address this issue, this article proposes a cascade perspective network (CPNet) to uncover reliable disagreement for subjective tasks. Specifically, CPNet learns each annotator’s personalized knowledge from annotation disagreement and stores them in an annotator bank through the personal perspective module for abundant disagreement information. Then, CPNet obtains consistent opinions by referring to all annotators’ opinions from the annotator bank through the comprehensive perspective module to reduce bias caused by noise. CPNet improves decision-making by considering the diversity and comprehensiveness of all annotators’ opinions. Moreover, it performs well in subjective tasks with limited or numerous annotators. The state-of-the-art (SOTA) results on subjective datasets from different domains demonstrate the effectiveness and generalizability of CPNet. Tong Zhang 0015, Canhui Zhang, Shuzhen Li, C. L. Philip Chen |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2025 | Cyclic Data Distillation Semi-Supervised Learning for Multi-Modal Emotion RecognitionabstractMulti-modal emotion recognition (MER) integrates multi-modal signals to help computers comprehensively understand human emotions, which is a crucial technology in human-computer interactions. However, the amount of labeled multi-modal emotion data is small and limits MER performance due to its expensive manual annotations. Meanwhile, semi-supervised learning (SSL) methods improving MER models with enormous unlabeled data suffer from confirmation bias, resulting in biased data distribution. To tackle these challenges, this paper proposes a cyclic data distillation semi-supervised learning (CDD-SSL) for MER tasks. CDD-SSL leverages multiple pre-trained unimodal teacher models and confidence-boosting pseudo-labelling (CBPL) to boost the confidence of multi-modal ensemble outputs and distill reliable and class-representative data from numerous unlabeled data. It then utilizes reliable and less-biased data to train a multi-modal student model and provides feedback to update all unimodal teacher models. CDD-SSL is a cyclic teacher-student framework with a feedback mechanism that gradually mitigates confirmation bias and obtains an effective MER model. Experimental results on four benchmark datasets demonstrate that CDD-SSL achieves superior performance over both the semi-supervised methods and the state-of-the-art fully-supervised models in MER tasks. Shuzhen Li, Tong Zhang 0015, C. L. Philip Chen |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2024 | Snapshot Prompt Ensemble for Parameter-Efficient Soft Prompt TransferabstractSoft Prompt Transfer(SPT) uses well-trained soft prompts as initialization to improve prompt tuning efficiency. However, most methods in SPT learn only a single and task-specific prompt for each source task. It may not be suitable for the target task and results in poor transferability on target task. To address this issue, we propose Snapshot Prompt Ensemble (SPE) method for parameter-efficient soft prompt transfer. Specifically, SPE extracts multiple soft prompts from each source task via taking snapshots at different training phases of prompt tuning. SPE then adaptively ensembles the multiple soft prompts and obtains a fused and instance-dependent prompt for the target task by a cross-task attention module. SPE can effectively exploit the prompts at different training phases and provide a more suitable starting point for target prompt training. Extensive experiments on various NLU tasks demonstrate that SPE outperforms state-of-the-art methods while SPE tunes less than 0.4% of the parameters compared to full fine-tuning. C. L. Philip Chen, Shuzhen Li, Tony Zhang |
ICASSP | 3 |
| 2024 | Disentanglement Network: Disentangle the Emotional Features from Acoustic Features for Speech Emotion RecognitionabstractSpeech emotion recognition plays a crucial role in human-computer interaction. However, data distribution of speech signals varies among individuals for emotion recognition. It may guide models to focus more on identity information rather than emotional information, which impairs the generalization ability of models. To address this issue, this paper proposes a novel Disentanglement Network (DTNet) to disentangle emotional features from acoustic features. Specifically, DTNet first captures hidden identity features from acoustic features through an identity-aware module. Then, we design a disentanglement module to disentangle emotional features from acoustic features within the constraints of a reconstruction module and the hidden identity features. These modules enable the DTNet to extract more discriminative emotional features for emotion recognition. Experimental results on both speaker-independent and speaker-dependent settings have proven the effectiveness of DTNet, and this method achieves an unweighted accuracy (UA) of 74.8% on the IEMOCAP dataset and UA of 95.5% on the Emo-DB dataset, outperforming the state-of-the-art methods on both datasets. Zhichen Yuan, C. L. Philip Chen, Shuzhen Li, Tong Zhang 0015 |
ICASSP | 3 |
| 2024 | DrM: Mastering Visual Reinforcement Learning through Dormant Ratio MinimizationabstractVisual reinforcement learning (RL) has shown promise in continuous control tasks.
Despite its progress, current algorithms are still unsatisfactory in virtually every aspect of the performance such as sample efficiency, asymptotic performance, and their robustness to the choice of random seeds.
In this paper, we identify a major shortcoming in existing visual RL methods that is the agents often exhibit sustained inactivity during early training, thereby limiting their ability to explore effectively.
Expanding upon this crucial observation, we additionally unveil a significant correlation between the agents' inclination towards motorically inactive exploration and the absence of neuronal activity within their policy networks.
To quantify this inactivity, we adopt dormant ratio as a metric to measure inactivity in the RL agent's network.
Empirically, we also recognize that the dormant ratio can act as a standalone indicator of an agent's activity level, regardless of the received reward signals.
Leveraging the aforementioned insights, we introduce DrM, a method that uses three core mechanisms to guide agents' exploration-exploitation trade-offs by actively minimizing the dormant ratio.
Experiments demonstrate that DrM achieves significant improvements in sample efficiency and asymptotic performance with no broken seeds (76 seeds in total) across three continuous control benchmark environments, including DeepMind Control Suite, MetaWorld, and Adroit.
Most importantly, DrM is the first model-free algorithm that consistently solves tasks in both the Dog and Manipulator domains from the DeepMind Control Suite as well as three dexterous hand manipulation tasks without demonstrations in Adroit, all based on pixel observations. Guowei Xu 0001, Ruijie Zheng, Yongyuan Liang, Zhecheng Yuan, Tianying Ji, Yu Luo 0021, Xiaoyu Liu 0003, Pu Hua, Shuzhen Li, Yanjie Ze, Hal Daumé III, Furong Huang, Huazhe Xu |
ICLR | 11 |
| 2024 | Multi-Scale Prompt Memory-Augmented Model for Black-Box ScenariosabstractXiaojun Kuang, C. L. Philip Chen, Shuzhen Li, Tong Zhang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Xiaojun Kuang, C. L. Philip Chen, Shuzhen Li, Tong Zhang 0015 |
NAACL-HLT | 3 |
| 2024 | Multi-Granularity Temporal-Spectral Representation Learning for Speech Emotion RecognitionabstractSpeech emotion recognition (SER) captures emotional information from speech signals to recognize users' emotional states, which plays a crucial role in conversational human-computer interaction. Most SER researches focus on ex-ploiting emotional information from global temporal or spectral features, but it may neglect detailed emotion-related information such as phonemes and syllables. To address this problem, this paper proposes a multi-granularity temporal-spectral representation learning (MG-TSRL) network for speech emotion recognition tasks. Specifically, MG-TSRL extracts different temporal features in phonetic, syllabic, and sentential granular-ity from spectrograms to retain more detailed emotional-related information. It then designs multilayer emotion-aware units to capture emotion-related frequency patterns and obtain deep spectrum features at each temporal granularity feature. MG-TSRL further introduces a fast broad learning system and feeds deep temporal-spectral features to it to obtain more accurate emotions. MG-TSRL gradually achieves effective temporal-spectral representation learning through multi-granularity temporal features and multilayer frequency pattern learning. The state-of-the-art results on the CASIA, RAVDESS, and SAVEE datasets are respectively 95.17%, 92.78%, and 87.50% in unweighted accuracy, demonstrating the effectiveness of MG-TSRL in speech emotion recognition. Zhichen Yuan, C. L. Philip Chen, Shuzhen Li, Tong Zhang 0015 |
SMC | 3 |
| 2024 | Deuce: Dual-diversity Enhancement and Uncertainty-awareness for Cold-start Active LearningabstractAbstract Cold-start active learning (CSAL) selects valuable instances from an unlabeled dataset for manual annotation. It provides high-quality data at a low annotation cost for label-scarce text classification. However, existing CSAL methods overlook weak classes and hard representative examples, resulting in biased learning. To address these issues, this paper proposes a novel dual-diversity enhancing and uncertainty-aware (Deuce) framework for CSAL. Specifically, Deuce leverages a pretrained language model (PLM) to efficiently extract textual representations, class predictions, and predictive uncertainty. Then, it constructs a Dual-Neighbor Graph (DNG) to combine information on both textual diversity and class diversity, ensuring a balanced data distribution. It further propagates uncertainty information via density-based clustering to select hard representative instances. Deuce performs well in selecting class-balanced and hard representative data by dual-diversity and informativeness. Experiments on six NLP datasets demonstrate the superiority and efficiency of Deuce. C. L. Philip Chen, Shuzhen Li, Tong Zhang 0015 |
Trans. Assoc. Comput. Linguistics | 3 |
| 2024 | SIA-Net: Sparse Interactive Attention Network for Multimodal Emotion RecognitionabstractMultimodal emotion recognition (MER) integrates multiple modalities to identify the user's emotional state, which is the core technology of natural and friendly human–computer interaction systems. Currently, many researchers have explored comprehensive multimodal information for MER, but few consider that comprehensive multimodal features may contain noisy, useless, or redundant information, which interferes with emotional feature representation. To tackle this challenge, this article proposes a sparse interactive attention network (SIA-Net) for MER. In SIA-Net, the sparse interactive attention (SIA) module mainly consists of intramodal sparsity and intermodal sparsity. The intramodal sparsity provides sparse but effective unimodal features for multimodal fusion. The intermodal sparsity adaptively sparses intramodal and intermodal interactive relations and encodes them into sparse interactive attention. The sparse interactive attention with a small number of nonzero weights then act on multimodal features to highlight a few but important features and suppress numerous redundant features. Furthermore, the intramodal sparsity and intermodal sparsity are deep sparse representations that make unimodal features and multimodal interactions sparse without complicated optimization. The extensive experimental results show that SIA-Net achieves superior performance on three widely used datasets. Shuzhen Li, Tong Zhang 0015, C. L. Philip Chen |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2024 | TT-GCN: Temporal-Tightly Graph Convolutional Network for Emotion Recognition From GaitsabstractThe human gait reflects substantial information about individual emotions. Current gait emotion recognition methods focus on capturing gait topology information and ignore the importance of fine-grained temporal features. This article proposes the temporal-tightly graph convolutional network (TT-GCN) to extract temporal features. TT-GCN comprises three significant mechanisms: the causal temporal convolution network (casual-TCN), the walking direction recognition auxiliary task, and the feature mapping layer. To obtain tight temporal dependencies and enhance the relevance among gait periods, the causal-TCN is introduced. Based on the assumption of emotional consistency in the walking directions, the auxiliary task is proposed to enhance the ability of fine-grained feature extraction. Through the feature mapping layer, affective features can be mapped into the appropriate representation and fused with deep learning features. TT-GCN shows the best performance across five comprehensive metrics. All experimental results verify the necessity and feasibility of exploring fine-grained temporal feature extraction. Tong Zhang 0015, Yelin Chen, Shuzhen Li, Xiping Hu, C. L. Philip Chen |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2023 | MIA-Net: Multi-Modal Interactive Attention Network for Multi-Modal Affective AnalysisabstractWhen a multi-modal affective analysis model generalizes from a bimodal task to a trimodal or multi-modal task, it is usually transformed into a hierarchical fusion model based on every two pairwise modalities, similar to a binary tree structure. This easily leads to large growth in model parameters and computation as the number of modalities increases, which limits the model's generalization. Moreover, many multi-modal fusion methods ignore that different modalities contribute differently to affective analysis. To tackle these challenges, this article proposes a general multi-modal fusion model that supports trimodal or multi-modal affective analysis tasks, called Multi-modal Interactive Attention Network (MIA-Net). Instead of treating different modalities equally, MIA-Net takes the modality that contributes the most to emotion as the main modality and the others as auxiliary modalities. MIA-Net introduces multi-modal interactive attention modules to adaptively select the important information of each auxiliary modality one by one to improve the main-modal representation. Moreover, MIA-Net enables quick generalization to trimodal or multi-modal tasks through stacking multiple MIA modules, which maintains efficient training and only requires linear computation and stable parameter counts. Experimental results of the transfer, generalization, and efficiency experiments on the widely-used datasets demonstrate the effectiveness and generalization of the proposed method. Shuzhen Li, Tong Zhang 0015, Bianna Chen, C. L. Philip Chen |
IEEE Trans. Affect. Comput. | 1 |
| 2023 | AIA-Net: Adaptive Interactive Attention Network for Text-Audio Emotion RecognitionabstractEmotion recognition based on text-audio modalities is the core technology for transforming a graphical user interface into a voice user interface, and it plays a vital role in natural human-computer interaction systems. Currently, mainstream multimodal learning research has designed various fusion strategies to learn intermodality interactions but hardly considers that not all modalities play equal roles in emotion recognition. Therefore, the main challenge in multimodal emotion recognition is how to implement effective fusion algorithms based on the auxiliary structure. To address this problem, this article proposes an adaptive interactive attention network (AIA-Net). In AIA-Net, text is treated as a primary modality, and audio is an auxiliary modality. AIA-Net adapts to textual and acoustic features with different dimensions and learns their dynamic interactive relations in a more flexible way. The interactive relations are encoded as interactive attention weights to focus on the acoustic features that are effective for textual emotional representations. AIA-Net performs well in adaptively assisting the textual emotional representation with the acoustic emotional information. Moreover, multiple collaborative learning (co-learning) layers of AIA-Net achieve multiple multimodal interactions and the deep bottom-up evolution of emotional representations. Experimental results on three benchmark datasets demonstrate the great effectiveness of the proposed method over the state-of-the-art methods. Tong Zhang 0015, Shuzhen Li, Bianna Chen, Haozhang Yuan, C. L. Philip Chen |
IEEE Trans. Cybern. | 2 |
| 2021 | Spatiotemporal and frequential cascaded attention networks for speech emotion recognition
Shuzhen Li, Xiaofen Xing, Weiquan Fan, Bolun Cai, Perry Fordson, Xiangmin Xu 0001 |
Neurocomputing | 1 |
| 2016 | Object proposals using SVM-based integrated modelabstractUtilizing object proposals as a preprocessing procedure has been shown its significance in many multimedia computing tasks. Most state-of-the-art methods devoted to finding a generic objectness measure for rating the possibilities of the initial sliding windows with or without objects. In fact, the object criteria vary from one objectness measure to another, which leads to the definite bottleneck for the single method. By observing the performance of the state-of-the-art in the large dataset, an integrated objectness model is proposed in this paper by accumulating the advantages from selected state-of-the-art techniques. First, the initial bounding boxes are generated by the strategy as same as the method with the highest object detection rate and slowest intersection over union drop. Second, these candidate boxes are re-scored based on each method's objectness system. Then, a score feature is obtained for each bounding box. A support vector machines (SVM) is utilized to train a general model on the training set constructed from a series of score vectors and the probabilistic scores for the testing boxes are predicted according to the learned model. The final proposals are ranked on account of the predicted scores. The evaluation on the challenging PASCAL VOC 2007 dataset shows that the proposed method has dominant concentration with better performance compared to the single state-of-the-art method. Wenjing Geng, Shuzhen Li, Tongwei Ren, Gangshan Wu |
IJCNN | 2 |
| 2015 | Saliency cuts based on adaptive triple thresholdingabstractSalient object detection attracts much attention for its effectiveness in numerous applications. However, how to effectively produce a high quality binary mask from a saliency map, named saliency cuts, is still an open problem. In this paper, we propose a novel saliency cuts approach using unsupervised seeds generation and GrabCut algorithm. With the input of a saliency map, we produce seeds for segmentation using adaptive triple thresholding, and feed the seeds to GrabCut algorithm. Finally, a high quality object mask is generated by iteratively optimization. The experimental results show that the proposed approach is competent to the task of saliency cuts and outperforms the state-of-the-art methods. Shuzhen Li, Ran Ju, Tongwei Ren, Gangshan Wu |
ICIP | 1 |