Zheng Lian 0004

dblp:213/7387-4 · DBLP profile ↗
← Back
47ranked-venue papers
18as first author
38since 2021 · last 2026
0000-0001-9477-0599ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 14 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 9 first-author · 17 since 2021
YearPublicationVenuePosition
2026 AStar: Boosting Multimodal Reasoning with Automated Structured Thinking
abstract
Multimodal large language models excel across diverse domains but struggle with complex visual reasoning tasks. To enhance their reasoning capabilities, current approaches typically rely on explicit search or post-training techniques. However, search-based methods suffer from computational inefficiency due to extensive solution space exploration, while post-training methods demand substantial data, computational resources, and often exhibit training instability. To address these challenges, we propose **AStar**, a training-free, **A**utomatic **S**tructured **t**hinking paradigm for multimod**a**l **r**easoning. Specifically, we introduce novel "thought cards", a lightweight library of high-level reasoning patterns abstracted from prior samples. For each test problem, AStar adaptively retrieves the optimal thought cards and seamlessly integrates these external explicit guidelines with the model’s internal implicit reasoning capabilities. Compared to previous methods, AStar eliminates computationally expensive explicit search and avoids additional complex post-training processes, enabling a more efficient reasoning approach. Extensive experiments demonstrate that our framework achieves 53.9% accuracy on MathVerse (surpassing GPT-4o's 50.2%) and 32.7% on MathVision (outperforming GPT-4o's 30.4%). Further analysis reveals the remarkable transferability of our method: thought cards generated from mathematical reasoning can also be applied to other reasoning tasks, even benefiting general visual perception and understanding. AStar serves as a plug-and-play test-time inference method, compatible with other post-training techniques, providing an important complement to existing multimodal reasoning approaches.
Mingkuan Feng, Guocheng Zhai, Shuai Zhang 0014, Zheng Lian 0004, Fangrui Lv, Pengpeng Shao, Ruihan Jin, Zhengqi Wen, Jianhua Tao 0001
AAAI5
2026 Beyond Examples: Towards Automated Thought-level In-Context Reasoning for Large Language Models
abstract
Jinyang Wu, Mingkuan Feng, Shuai Zhang, Feihu Che, Zhengqi Wen, Chonghua Liao, Ling Yang, Haoran Luo, Zheng Lian, Jianhua Tao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Mingkuan Feng, Shuai Zhang 0014, Feihu Che, Zhengqi Wen, Chonghua Liao, Zheng Lian 0004, Jianhua Tao 0001
ACL (1)9
2026 MERBench: A Unified Evaluation Benchmark for Multimodal Emotion Recognition
abstract
Multimodal emotion recognition plays a vital role in enhancing user experience in human-computer interaction. Over the past few decades, researchers have developed a range of algorithms and made remarkable progress. While each approach demonstrates certain advantages, inconsistent choices in feature extraction methods, evaluation protocols, and experimental settings have hindered fair comparisons among them. These inconsistencies significantly impede the advancement of the field. To address this issue, we introduce MERBench, a unified evaluation benchmark for multimodal emotion recognition. Our goal is to assess the contributions of several key techniques commonly used in prior studies, such as feature selection, multimodal fusion, robustness analysis, fine-tuning, and pre-training. We believe this work offers clear and comprehensive guidance for future research. Based on the evaluation results of MERBench, we further point out some promising research directions. In addition, we present a new emotion dataset, MER2023, specifically designed for the Chinese language environment. This dataset serves as a benchmark for research in multi-label learning, noise robustness, and semi-supervised learning.
Zheng Lian 0004, Licai Sun, Yong Ren 0006, Haiyang Sun 0004, Lan Chen 0005, Bin Liu 0041, Jianhua Tao 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 IRNet: Iterative Refinement Network for Noisy Partial Label Learning
abstract
Partial label learning (PLL) is a typical weakly supervised learning, where each sample is associated with a set of candidate labels. Its basic assumption is that the ground-truth label must be in the candidate set, but this assumption may not be satisfied due to the unprofessional judgment of annotators. Therefore, we relax this assumption and focus on a more general task, noisy PLL, where the ground-truth label may not exist in the candidate set. To address this challenging task, we propose a novel framework called "Iterative Refinement Network (IRNet)", aiming to purify noisy samples through two key modules (i.e., noisy sample detection and label correction). To achieve better performance, we exploit smoothness constraints to reduce prediction errors in these modules. Through theoretical analysis, we prove that IRNet is able to reduce the noise level of the dataset and eventually approximate the Bayes optimal classifier. Meanwhile, IRNet is a plug-in strategy that can be integrated with existing PLL approaches. Experimental results on multiple benchmark datasets show that IRNet outperforms state-of-the-art approaches on noisy PLL.
Zheng Lian 0004, Lan Chen 0005, Licai Sun, Bin Liu 0041, Lei Feng 0006, Jianhua Tao 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 CMDPAD: A Chinese multimodal dynamic personality and affect dataset for affect prediction in conversations
Zisen Zhou, Chang Wen, Xuefei Liu, Jianhua Tao 0001, Zhengqi Wen, Zheng Lian 0004, Jinming Zhao, Bingsen Xiong, Shaozheng Qin
Pattern Recognit.8
2025 MEIJU - The 1st Multimodal Emotion and Intent Joint Understanding Challenge
abstract
Multimodal Emotion and Intent Joint Understanding (MEIJU) aims to decode the semantic information expressed in the multimodal dialogues while inferring the emotions and intents, providing users with a more humanized human-machine interaction experience. However, challenges such as difficulties in data acquisition and imbalance annotations have made it difficult for current methods to meet the demands of practical applications. Therefore, we have organized two tracks focusing on the key themes of semi-supervised learning and class imbalance. Additionally, we have prepared data in two different languages (English and Mandarin) for each track, treating each language as a sub-track, to encourage participants to explore solutions in more diverse linguistic environments. Our code can be found at https://github.com/AI-S2-Lab/MEIJU2025-baseline.
Rui Liu 0008, Xiaofen Xing, Zheng Lian 0004, Haizhou Li 0001, Björn W. Schuller, Haolin Zuo
ICASSP3
2025 Adversarial Training and Gradient Optimization for Partially Deepfake Audio Localization
abstract
Partially deepfake audio localization is important in audio forensics. However, existing localization models for partially deepfake audio face two major challenges: distribution shifts between training and testing data as well as insufficient utilization of information from both manipulated regions and boundaries. To address these challenges, we propose to use Adversarial training and Gradient Optimization (AGO) to improve partially fake audio localization. Specifically, we apply a gradient reversal layer to reduce the dependence on domain-specific features, enhancing the model’s generalization ability. Additionally, we introduce an alternating update strategy to learn information from both manipulated regions and boundaries, while orthogonal gradient updates minimize conflicts between the two tasks. We evaluated AGO on both the ADD2023 track 2 and PartialSpoof datasets. We achieved a 22.82% relative improvement over the first-ranked method of the ADD2023 track 2. We also achieved state-of-the-art results on the PartialSpoof dataset. Our code is available at https://github.com/Little-dingding/ATGO.
Siding Zeng, Jiangyan Yi, Jianhua Tao 0001, Zheng Lian 0004, Shan Liang 0007, Chuyuan Zhang, Yujie Chen 0006, Xiaohui Zhang 0006
ICASSP5
2025 AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models
abstract
The emergence of multimodal large language models (MLLMs) advances multimodal emotion recognition (MER) to the next level—from naive discriminative tasks to complex emotion understanding with advanced video understanding abilities and natural language description. However, the current community suffers from a lack of large-scale datasets with intensive, descriptive emotion annotations, as well as a multimodal-centric framework to maximize the potential of MLLMs for emotion understanding. To address this, we establish a new benchmark for MLLM-based emotion understanding with a novel dataset (MER-Caption) and a new model (AffectGPT). Utilizing our model-based crowd-sourcing data collection strategy, we construct the largest descriptive emotion dataset to date (by far), featuring over 2K fine-grained emotion categories across 115K samples. We also introduce the AffectGPT model, designed with pre-fusion operations to enhance multimodal integration. Finally, we present MER-UniBench, a unified benchmark with evaluation metrics tailored for typical MER tasks and the free-form, natural language output style of MLLMs. Extensive experimental results show AffectGPT's robust performance across various MER tasks. We have released both the code and the dataset to advance research and development in emotion understanding: https://github.com/zeroQiaoba/AffectGPT.
Zheng Lian 0004, Haoyu Chen 0001, Lan Chen 0005, Haiyang Sun 0004, Licai Sun, Yong Ren 0006, Zebang Cheng, Bin Liu 0041, Rui Liu 0008, Xiaojiang Peng, Jiangyan Yi, Jianhua Tao 0001
ICML1
2025 OV-MER: Towards Open-Vocabulary Multimodal Emotion Recognition
abstract
Multimodal Emotion Recognition (MER) is a critical research area that seeks to decode human emotions from diverse data modalities. However, existing machine learning methods predominantly rely on predefined emotion taxonomies, which fail to capture the inherent complexity, subtlety, and multi-appraisal nature of human emotional experiences, as demonstrated by studies in psychology and cognitive science. To overcome this limitation, we advocate for introducing the concept of open vocabulary into MER. This paradigm shift aims to enable models to predict emotions beyond a fixed label space, accommodating a flexible set of categories to better reflect the nuanced spectrum of human emotions. To achieve this, we propose a novel paradigm: Open-Vocabulary MER (OV-MER), which enables emotion prediction without being confined to predefined spaces. However, constructing a dataset that encompasses the full range of emotions for OV-MER is practically infeasible; hence, we present a comprehensive solution including a newly curated database, novel evaluation metrics, and a preliminary benchmark. By advancing MER from basic emotions to more nuanced and diverse emotional states, we hope this work can inspire the next generation of MER, enhancing its generalizability and applicability in real-world scenarios. Code and dataset are available at: https://github.com/zeroQiaoba/AffectGPT.
Zheng Lian 0004, Haiyang Sun 0004, Licai Sun, Haoyu Chen 0001, Lan Chen 0005, Zhuofan Wen 0001, Hailiang Yao, Bin Liu 0041, Rui Liu 0008, Shan Liang 0007, Ya Li 0001, Jiangyan Yi, Jianhua Tao 0001
ICML1
2025 MER 2025: When Affective Computing Meets Large Language Models
abstract
MER2025 is the third year of our MER series of challenges. Previously, MER2023 (http://merchallenge.cn/mer2023) focused on multi-label learning, noise robustness, and semi-supervised learning, while MER2024 (https://zeroqiaoba.github.io/MER2024-website) introduced a new track dedicated to open-vocabulary emotion recognition. This year, MER2025 centers on the theme ''When Affective Computing Meets Large Language Models (LLMs)''. We aim to shift the paradigm from traditional categorical frameworks reliant on predefined emotion taxonomies to LLM-driven generative methods, offering innovative solutions for more accurate and reliable emotion understanding. The challenge contains four tracks: MER-SEMI focuses on fixed categorical emotion recognition enhanced by semi-supervised learning; MER-FG explores fine-grained emotions, expanding recognition from basic to nuanced emotional states; MER-DES incorporates multimodal cues (beyond emotion words) into predictions to enhance model interpretability; MER-PR reveals whether emotion prediction results can improve personality recognition performance. For the first three tracks, the baseline code is available at MERTools (https://github.com/zeroQiaoba/MERTools) and datasets can be accessed via Hugging Face (https://huggingface.co/datasets/MERChallenge/MER2025). For the last track, the dataset and baseline code are available on GitHub (https://github.com/cai-cong/MER25_personality).
Zheng Lian 0004, Rui Liu 0008, Kele Xu, Bin Liu 0041, Xuefei Liu, Yazhou Zhang 0001, Xin Liu 0012, Yong Li 0032, Zebang Cheng, Haolin Zuo, Ziyang Ma 0001, Xiaojiang Peng, Xie Chen 0001, Ya Li 0001, Erik Cambria, Guoying Zhao 0001, Björn W. Schuller, Jianhua Tao 0001
ACM Multimedia1
2025 MRAC 2025: 3rd International Workshop on Multimodal, Generative and Responsible Affective Computing
abstract
Multimodal, generative, and responsible affective computing aims to enhance people's lives. In recent years, the AI revolution has already begun to impact daily life, with virtual assistants being deployed across various sectors such as healthcare, banking, transportation, and education. It is clear that, in the near future, humans may interact with AI-powered systems as much or maybe even more than direct human-to-human interactions. Affective computing has numerous applications, including innovative approaches to forecasting and preventing anxiety, stress, and mental health issues; enhancing robotic empathy; assisting individuals with communication, behavior, and emotion regulation challenges; and promoting awareness of health and well-being. Many of these applications require enhanced control and protection of sensitive, private, and personal data. Therefore, it is crucial to further develop the creation, evaluation, and deployment of emotionally intelligent systems that are both responsive and responsible. Additionally, improving the accuracy and interpretability of emotion prediction results can significantly enhance the application of this technology in the downstream tasks mentioned above. MRAC'25 is the continuation of MRAC'23 and MRAC'24. Through this workshop, we aim to bring together researchers to discuss the potential and development of affective computing.
Zheng Lian 0004, Shreya Ghosh 0001, Erik Cambria, Zhixi Cai, Guoying Zhao 0001, Abhinav Dhall, Björn W. Schuller, Roland Göcke, Jianhua Tao 0001, Tom Gedeon
ACM Multimedia1
2025 Hardness-Aware Dynamic Curriculum Learning for Robust Multimodal Emotion Recognition with Missing Modalities
abstract
Missing modalities have recently emerged as a critical research direction in multimodal emotion recognition (MER). Conventional approaches typically address this issue through missing modality reconstruction. However, these methods fail to account for variations in reconstruction difficulty across different samples, consequently limiting the model's ability to handle hard samples effectively. To overcome this limitation, we propose a novel Hardness-Aware Dynamic Curriculum Learning framework, termed HARDY-MER. Our framework operates in two key stages: first, it estimates the hardness level of each sample, and second, it strategically emphasizes hard samples during training to enhance model performance on these challenging instances. Specifically, we first introduce a Multi-view Hardness Evaluation mechanism that quantifies reconstruction difficulty by considering both Direct Hardness (modality reconstruction errors) and Indirect Hardness (cross-modal mutual information). Meanwhile, we introduce a Retrieval-based Dynamic Curriculum Learning strategy that dynamically adjusts the training curriculum by retrieving samples with similar semantic information and balancing the learning focus between easy and hard instances. Extensive experiments on benchmark datasets demonstrate that HARDY-MER consistently outperforms existing methods in missing-modality scenarios. Our code will be made publicly available at https://github.com/HARDY-MER/HARDY-MER.
Rui Liu 0008, Haolin Zuo, Zheng Lian 0004, Hongyu Yuan
ACM Multimedia3
2025 ALLM4ADD: Unlocking the Capabilities of Audio Large Language Models for Audio Deepfake Detection
abstract
Audio deepfake detection (ADD) has grown increasingly important due to the rise of high-fidelity audio generative models and their potential for misuse. Given that audio large language models (ALLMs) have made significant progress in various audio processing tasks, a heuristic question arises: Can ALLMs be leveraged to solve ADD?. In this paper, we first conduct a comprehensive zero-shot evaluation of ALLMs on ADD, revealing their ineffectiveness. To this end, we propose ALLM4ADD, an ALLM-driven framework for ADD. Specifically, we reformulate ADD task as an audio question answering problem, prompting the model with the question: ''Is this audio fake or real?''. We then perform supervised fine-tuning to enable the ALLM to assess the authenticity of query audio. Extensive experiments are conducted to demonstrate that our ALLM-based method can achieve superior performance in fake audio detection, particularly in data-scarce scenarios. As a pioneering study, we anticipate that this work will inspire the research community to leverage ALLMs to develop more effective ADD systems. Code is available at https://github.com/ucas-hao/qwen_audio_for_add.git.
Jiangyan Yi, Chenglong Wang 0001, Jianhua Tao 0001, Zheng Lian 0004, Yong Ren 0006, Yujie Chen 0006, Zhengqi Wen
ACM Multimedia5
2025 Enhancing Multimodal Personality Assessment with LLM-Augmented Hierarchical Fusion
abstract
This study proposes an LLM-augmented hierarchical fusion framework to enhance multimodal personality and ability assessment for the ACM MULTIMEDIA AVI CHALLENGE 2025, addressing semantic sparsity and cross-modal interaction limitations. We leverage large language models (e.g., Qwen, DeepSeek) to generate psychologically enriched text descriptions, bridging raw transcripts with expert evaluations, and integrate them with audio-visual features through early fusion and multi-path MLP ensembles. Track 1 (personality regression) employs dual-text inputs while Track 2 (multi-label ability prediction) uses parallel regression. Results show significant improvements: 23.1% MSE reduction over text-only baselines in Track 1, and 10.7%/12.5% gains over state-of-the-art fusion in Tracks 1/2, with 31.2% average improvement for cognitive traits (Q3-Q5). The framework demonstrates the effectiveness of semantic enhancement and adaptive fusion, with future work focusing on overfitting mitigation and feature optimization.
Longjiang Yang, Zhuofan Wen 0001, Hailiang Yao, Bin Liu 0041, Zheng Lian 0004, Jianhua Tao 0001
ACM Multimedia10
2025 MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
abstract
We introduce MMAR, a new benchmark designed to evaluate the deep reasoning capabilities of Audio-Language Models (ALMs) across massive multi-disciplinary tasks. MMAR comprises 1,000 meticulously curated audio-question-answer triplets, collected from real-world internet videos and refined through iterative error corrections and quality checks to ensure high quality. Unlike existing benchmarks that are limited to specific domains of sound, music, or speech, MMAR extends them to a broad spectrum of real-world audio scenarios, including mixed-modality combinations of sound, music, and speech. Each question in MMAR is hierarchically categorized across four reasoning layers: Signal, Perception, Semantic, and Cultural, with additional sub-categories within each layer to reflect task diversity and complexity. To further foster research in this area, we annotate every question with a Chain-of-Thought (CoT) rationale to promote future advancements in audio reasoning. Each item in the benchmark demands multi-step deep reasoning beyond surface-level understanding. Moreover, a part of the questions requires graduate-level perceptual and domain-specific knowledge, elevating the benchmark's difficulty and depth. We evaluate MMAR using a broad set of models, including Large Audio-Language Models (LALMs), Large Audio Reasoning Models (LARMs), Omni Language Models (OLMs), Large Language Models (LLMs), and Large Reasoning Models (LRMs), with audio caption inputs. The performance of these models on MMAR highlights the benchmark's challenging nature, and our analysis further reveals critical limitations of understanding and reasoning capabilities among current models. These findings underscore the urgent need for greater research attention in audio-language reasoning, including both data and algorithm innovation. We hope MMAR will serve as a catalyst for future advances in this important but little-explored area.
Ziyang Ma 0001, Yinghao Ma, Yanqiao Zhu 0003, Yi-Wen Chao, Yuanzhe Chen, Zhuo Chen 0006, Jian Cong, Keliang Li, Siyou Li, Xinfeng Li, Xiquan Li, Zheng Lian 0004, Yuzhe Liang, Minghao Liu 0003, Zhikang Niu, Tianrui Wang, Yuping Wang 0005, Yuxuan Wang 0002, Guanrou Yang, Jianwei Yu 0001, Ruibin Yuan, Zhisheng Zheng, Ziya Zhou, Haina Zhu, Wei Xue 0002, Emmanouil Benetos, Kai Yu 0004, Chng Eng Siong, Xie Chen 0001
NeurIPS16
2025 Are MLLMs Trapped in the Visual Room?
Yazhou Zhang 0001, Chunwang Zou, Qimeng Liu, Lu Rong, Ben Yao, Zheng Lian 0004, Qiuchi Li, Peng Zhang 0002, Harry Qin
PRCV (7)6
2025 SVFAP: Self-Supervised Video Facial Affect Perceiver
abstract
Video-based facial affect analysis has recently attracted increasing attention owing to its critical role in human-computer interaction. Previous studies mainly focus on developing various deep learning architectures and training them in a fully supervised manner. Although significant progress has been achieved by these supervised methods, the longstanding lack of large-scale high-quality labeled data severely hinders their further improvements. Motivated by the recent success of self-supervised learning in computer vision, this paper introduces a self-supervised approach, termed Self-supervised Video Facial Affect Perceiver (SVFAP), to address the dilemma faced by supervised methods. Specifically, SVFAP leverages masked facial video autoencoding to perform self-supervised pre-training on massive unlabeled facial videos. Considering that large spatiotemporal redundancy exists in facial videos, we propose a novel temporal pyramid and spatial bottleneck Transformer as the encoder of SVFAP, which not only largely reduces computational costs but also achieves excellent performance. To verify the effectiveness of our method, we conduct experiments on nine datasets spanning three downstream tasks, including dynamic facial expression recognition, dimensional emotion recognition, and personality recognition. Comprehensive results demonstrate that SVFAP can learn powerful affect-related representations via large-scale self-supervised pre-training and it significantly outperforms previous state-of-the-art methods on all datasets.
Licai Sun, Zheng Lian 0004, Haiyang Sun 0004, Bin Liu 0041, Jianhua Tao 0001
IEEE Trans. Affect. Comput.2
2025 SarcasmBench: Towards Evaluating Large Language Models on Sarcasm Understanding
abstract
In the era of large language models (LLMs), tasks associated with “System I” cognition—those that are fast, automatic, and intuitive, such as sentiment analysis and text classification—are often considered effectively solved. However, sarcasm remains a persistent challenge. As a subtle and complex linguistic phenomenon, sarcasm frequently involves rhetorical devices such as hyperbole and figurative language to express implicit sentiments and intentions, demanding a higher level of abstraction and pragmatic reasoning than standard sentiment analysis. This raises concerns about whether current claims of LLM success extend robustly to the domain of sarcasm understanding. To systematically investigate this issue, we introduce a new high-quality multi-modal sarcasm detection dataset, termedAMSD, and construct a comprehensive evaluation benchmark,SarcasmBench. Our benchmark encompasses 16 state-of-the-art (SOTA) LLMs and 8 strong pretrained language models (PLMs), evaluated across six widely-used textual sarcasm datasets and three multi-modal sarcasm benchmarks. We adopt three popular prompting paradigms: zero-shot input/output (IO) prompting, few-shot IO prompting, and chain-of-thought (CoT) prompting. Our extensive experiments yield three key findings: (1) current LLMs underperform supervised PLMs based sarcasm detection baselines. This suggests that significant efforts are still required to improve LLMs' understanding of human sarcasm. (2) GPT-4 and Gemini 2.0 consistently and significantly outperforms other LLMs across various prompting methods. (3) Few-shot IO prompting method outperforms the other two methods: zero-shot IO and few-shot CoT. We hope this benchmark will serve as a valuable resource for the research community and inspire future work toward more robust and human-aligned sarcasm understanding.
Yazhou Zhang 0001, Chunwang Zou, Zheng Lian 0004, Prayag Tiwari, Harry Qin
IEEE Trans. Affect. Comput.3
2024 NLoPT: N-gram Enhanced Low-Rank Task Adaptive Pre-training for Efficient Language Model Adaption
abstract
Pre-trained Language Models (PLMs) like BERT have achieved superior performance on different downstream tasks, even when such a model is trained on a general domain. Moreover, recent studies have shown that continued pre-training on task-specific data, known as task adaptive pre-training (TAPT), can further improve downstream task performance. However, conventional TAPT adjusts all the parameters of the PLMs, which distorts the learned generic knowledge embedded in the original PLMs weights, and it is expensive to store a whole model copy for each downstream task. In this paper, we propose NLoPT, a two-step n-gram enhanced low-rank task adaptive pre-training method, to effectively and efficiently customize a PLM to the downstream task. Specifically, we first apply low-rank adaption (LoRA), a prevalent parameter-efficient technique, for efficient TAPT. We further explicitly incorporate the task-specific multi-granularity n-gram information via the cross-attention mechanism. Experimental results on six datasets from four domains illustrate the effectiveness of NLoPT, demonstrating the superiority of LoRA based TAPT and the necessity of incorporating task-specific n-gram information.
Jiangyan Yi, Zheng Lian 0004, Jianhua Tao 0001, Xinrui Yan
LREC/COLING3
2024 Pseudo Labels Regularization for Imbalanced Partial-Label Learning
abstract
Partial-label learning (PLL) is an important branch of weakly supervised learning where the single ground truth resides in a set of candidate labels, while the research rarely considers the label imbalance. A recent study for imbalanced PLL propose that the combinatorial challenge of partial-label learning and long-tail learning lies in matching between a decent marginal prior distribution with drawing the pseudo labels. However, even if the pseudo label matches the prior distribution, the tail classes will still be difficult to learn because the total weight of tail classes is too small. Therefore, we propose a pseudo-label regularization technique specially designed for imbalanced PLL. By punishing the pseudo labels of head classes, our method implements state-of-art under the standardized benchmarks compared to the previous PLL methods.
Zheng Lian 0004, Bin Liu 0041, Zerui Chen, Jianhua Tao 0001
ICASSP2
2024 MFSN: Multi-perspective Fusion Search Network For Pre-training Knowledge in Speech Emotion Recognition
Haiyang Sun 0004, Fulin Zhang, Yingying Gao, Shilei Zhang, Zheng Lian 0004, Junlan Feng
INTERSPEECH5
2024 Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning
abstract
Accurate emotion perception is crucial for various applications, including human-computer interaction, education, and counseling. However, traditional single-modality approaches often fail to capture the complexity of real-world emotional expressions, which are inherently multimodal. Moreover, existing Multimodal Large Language Models (MLLMs) face challenges in integrating audio and recognizing subtle facial micro-expressions. To address this, we introduce the MERR dataset, containing 28,618 coarse-grained and 4,487 fine-grained annotated samples across diverse emotional categories. This dataset enables models to learn from varied scenarios and generalize to real-world applications. Furthermore, we propose Emotion-LLaMA, a model that seamlessly integrates audio, visual, and textual inputs through emotion-specific encoders. By aligning features into a shared space and employing a modified LLaMA model with instruction tuning, Emotion-LLaMA significantly enhances both emotional recognition and reasoning capabilities. Extensive evaluations show Emotion-LLaMA outperforms other MLLMs, achieving top scores in Clue Overlap (7.83) and Label Overlap (6.25) on EMER, an F1 score of 0.9036 on MER2023-SEMI challenge, and the highest UAR (45.59) and WAR (59.37) in zero-shot evaluations on DFEW dataset.
Zebang Cheng, Zhi-Qi Cheng, Jun-Yan He, Kai Wang 0036, Zheng Lian 0004, Xiaojiang Peng, Alex Hauptmann 0001
NeurIPS6
2024 Contrastive Learning Based Modality-Invariant Feature Acquisition for Robust Multimodal Emotion Recognition With Missing Modalities
abstract
Multimodal emotion recognition (MER) aims to understand the way that humans express their emotions by exploring complementary information across modalities. However, it is hard to guarantee that full-modality data is always available in real-world scenarios. To deal with missing modalities, researchers focused on meaningful joint multimodal representation learning during cross-modal missing modality imagination. However, the cross-modal imagination mechanism is highly susceptible to errors due to the “modality gap” issue, which affects the imagination accuracy, thus, the final recognition performance. To this end, we introduce the concept of a modality-invariant feature into the missing modality imagination network, which contains two key modules: 1) a novel contrastive learning-based module to extract modality-invariant features under full modalities; 2) a robust imagination module based on imagined invariant features to reconstruct missing information under missing conditions. Finally, we incorporate imagined and available modalities for emotion recognition. Experimental results on benchmark datasets demonstrate that our proposed method outperforms existing state-of-the-art strategies. Compared with our previous work, our extended version is more effective on multimodal emotion recognition with missing modalities. The code is released athttps://github.com/ZhuoYulang/CIF-MMIN.
Rui Liu 0008, Haolin Zuo, Zheng Lian 0004, Björn W. Schuller, Haizhou Li 0001
IEEE Trans. Affect. Comput.3
2024 Efficient Multimodal Transformer With Dual-Level Feature Restoration for Robust Multimodal Sentiment Analysis
abstract
With the proliferation of user-generated online videos, Multimodal Sentiment Analysis (MSA) has attracted increasing attention recently. Despite significant progress, there are still two major challenges on the way towards robust MSA: 1) inefficiency when modeling cross-modal interactions in unaligned multimodal data; and 2) vulnerability to random modality feature missing which typically occurs in realistic settings. In this paper, we propose a generic and unified framework to address them, named Efficient Multimodal Transformer with Dual-Level Feature Restoration (EMT-DLFR). Concretely, EMT employs utterance-level representations from each modality as the global multimodal context to interact with local unimodal features and mutually promote each other. It not only avoids the quadratic scaling cost of previous local-local cross-modal interaction methods but also leads to better performance. To improve model robustness in the incomplete modality setting, on the one hand, DLFR performs low-level feature reconstruction to implicitly encourage the model to learn semantic information from incomplete data. On the other hand, it innovatively regards complete and incomplete data as two different views of one sample and utilizes siamese representation learning to explicitly attract their high-level representations. Comprehensive experiments on three popular datasets demonstrate that our method achieves superior performance in both complete and incomplete modality settings.
Licai Sun, Zheng Lian 0004, Bin Liu 0041, Jianhua Tao 0001
IEEE Trans. Affect. Comput.2
2024 PIRNet: Personality-Enhanced Iterative Refinement Network for Emotion Recognition in Conversation
abstract
Emotion recognition in conversation (ERC) is important for enhancing user experience in human-computer interaction. Unlike vanilla emotion recognition in individual utterances, ERC aims to classify constituent utterances in a dialog into corresponding emotion labels, which makes contextual information crucial. In addition to contextual information, personality traits also affect emotional perception based on psychological findings. Although researchers have proposed several approaches and achieved promising results on ERC, current works in this domain rarely incorporate contextual information and personality influence. To this end, we propose a novel framework to integrate these factors seamlessly, called "Personality-enhanced Iterative Refinement Network (PIRNet)." Specifically, PIRNet is a multistage iterative method. To capture personality influence, PIRNet leverages personality traits to mimic emotional transitions and generates personality-enhanced results. Then we exploit sequence models to capture contextual information in conversations. To verify the effectiveness of our proposed method, we conduct experiments on three benchmark datasets for ERC, that is, IEMOCAP, CMU-MOSI, and CMU-MOSEI. Experimental results demonstrate that our PIRNet succeeds over currently advanced approaches to emotion recognition.
Zheng Lian 0004, Bin Liu 0041, Jianhua Tao 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 EmotionNAS: Two-stream Neural Architecture Search for Speech Emotion Recognition
Haiyang Sun 0004, Zheng Lian 0004, Bin Liu 0041, Jianhua Tao 0001, Licai Sun, Cong Cai, Meng Wang 0001
INTERSPEECH2
2023 MRAC'23: 1st International Workshop on Multimodal and Responsible Affective Computing
abstract
Multimodal emotion recognition has become an important research topic due to its wide applications in human-computer interaction. Over the last few decades, the technology has made remarkable progress with the development of deep learning. However, existing technologies are hard to meet the demand for practical applications. To this end, we organize this workshop to bring together researchers in this field to further discuss recent research and future directions.
Zheng Lian 0004, Erik Cambria, Guoying Zhao 0001, Björn W. Schuller, Jianhua Tao 0001
ACM Multimedia1
2023 MER 2023: Multi-label Learning, Modality Robustness, and Semi-Supervised Learning
abstract
The first Multimodal Emotion Recognition Challenge (MER 2023)1 was successfully held at ACM Multimedia. The challenge focuses on system robustness and consists of three distinct tracks: (1) MER-MULTI, where participants are required to recognize both discrete and dimensional emotions; (2) MER-NOISE, in which noise is added to test videos for modality robustness evaluation; (3) MER-SEMI, which provides a large amount of unlabeled samples for semi-supervised learning. In this paper, we introduce the motivation behind this challenge, describe the benchmark dataset, and provide some statistics about participants. To continue using this dataset after MER 2023, please sign a new End User License Agreement2 and send it to our official email address3. We believe this high-quality dataset can become a new benchmark in multimodal emotion recognition, especially for the Chinese research community.
Zheng Lian 0004, Haiyang Sun 0004, Licai Sun, Jinming Zhao, Ye Liu 0010, Bin Liu 0041, Jiangyan Yi, Meng Wang 0001, Erik Cambria, Guoying Zhao 0001, Björn W. Schuller, Jianhua Tao 0001
ACM Multimedia1
2023 MAE-DFER: Efficient Masked Autoencoder for Self-supervised Dynamic Facial Expression Recognition
abstract
Dynamic facial expression recognition (DFER) is essential to the development of intelligent and empathetic machines. Prior efforts in this field mainly fall into supervised learning paradigm, which is severely restricted by the limited labeled data in existing datasets. Inspired by recent unprecedented success of masked autoencoders (e.g., VideoMAE), this paper proposes MAE-DFER, a novel self-supervised method which leverages large-scale self-supervised pre-training on abundant unlabeled data to largely advance the development of DFER. Since the vanilla Vision Transformer (ViT) employed in VideoMAE requires substantial computation during fine-tuning, MAE-DFER develops an efficient local-global interaction Transformer (LGI-Former) as the encoder. Moreover, in addition to the standalone appearance content reconstruction in VideoMAE, MAE-DFER also introduces explicit temporal facial motion modeling to encourage LGI-Former to excavate both static appearance and dynamic motion information. Extensive experiments on six datasets show that MAE-DFER consistently outperforms state-of-the-art supervised methods by significant margins (e.g., +6.30% UAR on DFEW and +8.34% UAR on MAFW), verifying that it can learn powerful dynamic facial representations via large-scale self-supervised pre-training. Besides, it has comparable or even better performance than VideoMAE, while largely reducing the computational cost (about 38% FLOPs). We believe MAE-DFER has paved a new way for the advancement of DFER and can inspire more relevant research in this field and even other related tasks. Codes and models are publicly available at https://github.com/sunlicai/MAE-DFER.
Licai Sun, Zheng Lian 0004, Bin Liu 0041, Jianhua Tao 0001
ACM Multimedia2
2023 Integrating VideoMAE based model and Optical Flow for Micro- and Macro-expression Spotting
abstract
The task of interval localization of macro- and micro-expression in long videos has a wide range of applications in the field of human-computer interaction. Compared with macro-expression, micro-expression has shorter duration, lower intensity, and smaller number of samples, which make them more difficult to spot accurately in long videos. In this paper, we propose a pre-trained model combined with the optical flow method to improve the accuracy and robustness of macro- and micro-expression spotting. Firstly, self-supervised pre-training is performed on rich unlabeled data based on VideoMAE. Then, multiple models are trained on the datasets SAMM-LV and CAS(ME)³ for macro- and micro-expression with different fine-grains. Finally, different lengths of slices are generated based on the models with different fine-grains, and the optimal matching method through the combination of model fine-grainedness and slice lengths is explored. At the same time, macro- and micro-expression generating regions were spotted using the optical flow method, fused with the model outputs to supplement the spatio-temporal information not captured by the model and to exclude the interference of non-interested regions. We evaluated the performance of our method on the MEGC2023 testset (consisting of 10 long videos from SAMM and 20 long videos from CAS(ME)3) and won first place in the MEGC2023 Challenge. The results demonstrate the effectiveness of the method.
Licai Sun, Zheng Lian 0004, Bin Liu 0041, Haiyang Sun 0004, Jianhua Tao 0001
ACM Multimedia4
2023 ALIM: Adjusting Label Importance Mechanism for Noisy Partial Label Learning
abstract
Noisy partial label learning (noisy PLL) is an important branch of weakly supervised learning. Unlike PLL where the ground-truth label must conceal in the candidate label set, noisy PLL relaxes this constraint and allows the ground-truth label may not be in the candidate label set. To address this challenging problem, most of the existing works attempt to detect noisy samples and estimate the ground-truth label for each noisy sample. However, detection errors are unavoidable. These errors can accumulate during training and continuously affect model optimization. To this end, we propose a novel framework for noisy PLL with theoretical interpretations, called ``Adjusting Label Importance Mechanism (ALIM)''. It aims to reduce the negative impact of detection errors by trading off the initial candidate set and model outputs. ALIM is a plug-in strategy that can be integrated with existing PLL approaches. Experimental results on multiple benchmark datasets demonstrate that our method can achieve state-of-the-art performance on noisy PLL. Our code is available at: https://github.com/zeroQiaoba/ALIM.
Zheng Lian 0004, Lei Feng 0006, Bin Liu 0041, Jianhua Tao 0001
NeurIPS2
2023 VRA: Variational Rectified Activation for Out-of-distribution Detection
abstract
Out-of-distribution (OOD) detection is critical to building reliable machine learning systems in the open world. Researchers have proposed various strategies to reduce model overconfidence on OOD data. Among them, ReAct is a typical and effective technique to deal with model overconfidence, which truncates high activations to increase the gap between in-distribution and OOD. Despite its promising results, is this technique the best choice? To answer this question, we leverage the variational method to find the optimal operation and verify the necessity of suppressing abnormally low and high activations and amplifying intermediate activations in OOD detection, rather than focusing only on high activations like ReAct. This motivates us to propose a novel technique called ``Variational Rectified Activation (VRA)'', which simulates these suppression and amplification operations using piecewise functions. Experimental results on multiple benchmark datasets demonstrate that our method outperforms existing post-hoc strategies. Meanwhile, VRA is compatible with different scoring functions and network architectures. Our code is available at https://github.com/zeroQiaoba/VRA.
Zheng Lian 0004, Bin Liu 0041, Jianhua Tao 0001
NeurIPS2
2023 GCNet: Graph Completion Network for Incomplete Multimodal Learning in Conversation
abstract
Conversations have become a critical data format on social media platforms. Understanding conversation from emotion, content and other aspects also attracts increasing attention from researchers due to its widespread application in human-computer interaction. In real-world environments, we often encounter the problem of incomplete modalities, which has become a core issue of conversation understanding. To address this problem, researchers propose various methods. However, existing approaches are mainly designed for individual utterances rather than conversational data, which cannot fully exploit temporal and speaker information in conversations. To this end, we propose a novel framework for incomplete multimodal learning in conversations, called "Graph Complete Network (GCNet)," filling the gap of existing works. Our GCNet contains two well-designed graph neural network-based modules, "Speaker GNN" and "Temporal GNN," to capture temporal and speaker dependencies. To make full use of complete and incomplete data, we jointly optimize classification and reconstruction tasks in an end-to-end manner. To verify the effectiveness of our method, we conduct experiments on three benchmark conversational datasets. Experimental results demonstrate that our GCNet is superior to existing state-of-the-art approaches in incomplete multimodal learning.
Zheng Lian 0004, Lan Chen 0005, Licai Sun, Bin Liu 0041, Jianhua Tao 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 SMIN: Semi-Supervised Multi-Modal Interaction Network for Conversational Emotion Recognition
abstract
Conversational emotion recognition is a crucial research topic in human-computer interactions. Due to the heavy annotation cost and inevitable label ambiguity, collecting large amounts of labeled data is challenging and expensive, which restricts the performance of current fully-supervised methods in this domain. To address this problem, researchers attempt to distill knowledge from unlabeled data via semi-supervised learning. However, most of these semi-supervised methods ignore multimodal interactive information, although recent works have proven that such interactive information is essential for emotion recognition. To this end, we propose a novel framework to seamlessly integrate semi-supervised learning with multimodal interactions, called “Semi-supervised Multi-modal Interaction Network (SMIN)”. SMIN contains two well-designed semi-supervised modules, “Intra-modal Interactive Module (IIM)” and “Cross-modal Interactive Module (CIM)” to learn intra- and cross-modal interactions. These two modules leverage additional unlabeled data to extract emotion-salient representations. To capture additional contextual information, we utilize the hierarchical recurrent networks followed with the hybrid fusion strategy to integrate multimodal features. These multimodal features are further utilized for conversational emotion recognition. Experimental results on four benchmark datasets (i.e., IEMOCAP, MELD, CMU-MOSI and CMU-MOSEI) demonstrate that SMIN succeeds over existing state-of-the-art strategies on emotion recognition.
Zheng Lian 0004, Bin Liu 0041, Jianhua Tao 0001
IEEE Trans. Affect. Comput.1
2023 Multimodal Spatiotemporal Representation for Automatic Depression Level Detection
abstract
Physiological studies have shown that there are some differences in speech and facial activities between depressive and healthy individuals. Based on this fact, we propose a novel spatio-temporal attention (STA) network and a multimodal attention feature fusion (MAFF) strategy to obtain the multimodal representation of depression cues for predicting the individual depression level. Specifically, we first divide the speech amplitude spectrum/video into fixed-length segments and input these segments into the STA network, which not only integrates the spatial and temporal information through attention mechanism, but also emphasizes the audio/video frames related to depression detection. The audio/video segment-level feature is obtained from the output of the last full connection layer of the STA network. Second, this article employs the eigen evolution pooling method to summarize the changes of each dimension of the audio/video segment-level features to aggregate them into the audio/video level feature. Third, the multimodal representation with modal complementary information is generated using the MAFF and inputs into the support vector regression predictor for estimating depression severity. Experimental results on the AVEC2013 and AVEC2014 depression databases illustrate the effectiveness of our method.
Mingyue Niu, Jianhua Tao 0001, Bin Liu 0041, Jian Huang 0014, Zheng Lian 0004
IEEE Trans. Affect. Comput.5
2021 Multimodal Cross- and Self-Attention Network for Speech Emotion Recognition
abstract
Speech Emotion Recognition (SER) requires a thorough understanding of both the linguistic content of an utterance (i.e., textual information) and how the speaker utters it (i.e., acoustic information). The one vital challenge in SER is how to effectively fuse these two kinds of information. In this paper, we propose a novel Multimodal Cross- and Self-Attention Network (MCSAN) to tackle this problem. The core of MCSAN is to employ the parallel cross- and self-attention modules to explicitly model both inter- and intra-modal interactions of audio and text. Specifically, the cross-attention module utilizes the cross-attention mechanism to guide one modality to attend to the other modality and update the features accordingly. Similarly, the self-attention module employs the self-attention mechanism to propagate information within each modality. We evaluate MCSAN on two benchmark datasets, IEMOCAP and MELD. Experimental results demonstrate that our proposed model achieves state-of-the-art performance on both datasets.
Licai Sun, Bin Liu 0041, Jianhua Tao 0001, Zheng Lian 0004
ICASSP4
2021 DECN: Dialogical emotion correction network for conversational emotion recognition
Zheng Lian 0004, Bin Liu 0041, Jianhua Tao 0001
Neurocomputing1
2021 CTNet: Conversational Transformer Network for Emotion Recognition
abstract
Emotion recognition in conversation is a crucial topic for its widespread applications in the field of human-computer interactions. Unlike vanilla emotion recognition of individual utterances, conversational emotion recognition requires modeling both context-sensitive and speaker-sensitive dependencies. Despite the promising results of recent works, they generally do not leverage advanced fusion techniques to generate the multimodal representations of an utterance. In this way, they have limitations in modeling the intra-modal and cross-modal interactions. In order to address these problems, we propose a multimodal learning framework for conversational emotion recognition, called conversational transformer network (CTNet). Specifically, we propose to use the transformer-based structure to model intra-modal and cross-modal interactions among multimodal features. Meanwhile, we utilize word-level lexical features and segment-level acoustic features as the inputs, thus enabling us to capture temporal information in the utterance. Additionally, to model context-sensitive and speaker-sensitive dependencies, we propose to use the multihead attention based bi-directional GRU component and speaker embeddings. Experimental results on the IEMOCAP and MELD datasets demonstrate the effectiveness of the proposed method. Our method shows an absolute 2.1~6.2% performance improvement on weighted average F1 over state-of-the-art strategies.
Zheng Lian 0004, Bin Liu 0041, Jianhua Tao 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2020 Multimodal Transformer Fusion for Continuous Emotion Recognition
abstract
Multimodal fusion increases the performance of emotion recognition because of the complementarity of different modalities. Compared with decision level and feature level fusion, model level fusion makes better use of the advantages of deep neural networks. In this work, we utilize the Transformer model to fuse audio-visual modalities on the model level. Specifically, the multi-head attention produces multimodal emotional intermediate representations from common semantic feature space after encoding audio and visual modalities. Meanwhile, it also can learn long-term temporal dependencies with self-attention mechanism effectively. The experiments, on the AVEC 2017 database, shows the superiority of model level fusion than other fusion strategies. Moreover, we combine the Transformer model and LSTM to further improve the performance, which achieves better results than other methods.
Jian Huang 0014, Jianhua Tao 0001, Bin Liu 0041, Zheng Lian 0004, Mingyue Niu
ICASSP4
2020 Learning Utterance-Level Representations with Label Smoothing for Speech Emotion Recognition
Jian Huang 0014, Jianhua Tao 0001, Bin Liu 0041, Zheng Lian 0004
INTERSPEECH4
2020 Context-Dependent Domain Adversarial Neural Network for Multimodal Emotion Recognition
Zheng Lian 0004, Jianhua Tao 0001, Bin Liu 0041, Jian Huang 0014, Zhanlei Yang, Rongjun Li
INTERSPEECH1
2020 Conversational Emotion Recognition Using Self-Attention Mechanisms and Graph Neural Networks
Zheng Lian 0004, Jianhua Tao 0001, Bin Liu 0041, Jian Huang 0014, Zhanlei Yang, Rongjun Li
INTERSPEECH1
2020 ARVC: An Auto-Regressive Voice Conversion System Without Parallel Training Data
Zheng Lian 0004, Zhengqi Wen, Xinyong Zhou, Songbai Pu, Shengkai Zhang, Jianhua Tao 0001
INTERSPEECH1
2019 Discriminative Video Representation with Temporal Order for Micro-expression Recognition
abstract
Micro-expression recognition is a challenging task due to its low intensity and short duration and how to extract the subtle facial changes is a key issue in this field. Although there are many methods attempt to cope with this problem, they are difficult to encode the temporal order of all frames in the video clips. For these reasons, this paper employs rank pooling and ℓ2,1-norm to obtain the discriminative video representation with temporal order. In particular, we extract Local Two-Order Gradient Pattern (LTOGP) feature of each frame to describe the subtle information. Then, the video representation is generated by using rank pooling, which captures the temporal order among all frames. Furthermore, considering the sparsity of ℓ2,1-norm, we can select those discriminant features. Finally, micro-expression classification is accomplished using SVM. Experiments are conducted on two publicly available micro-expression databases i.e. CASME and CASME2. The results demonstrate that our method achieves better performance than the state-of-the-art algorithms.
Mingyue Niu, Jianhua Tao 0001, Ya Li 0001, Jian Huang 0014, Zheng Lian 0004
ICASSP5
2019 Conversational Emotion Analysis via Attention Mechanisms
abstract
Different from the emotion recognition in individual utterances, we propose a multimodal learning framework using relation and dependencies among the utterances for conversational emotion analysis. The attention mechanism is applied to the fusion of the acoustic and lexical features. Then these fusion representations are fed into the self-attention based bi-directional gated recurrent unit (GRU) layer to capture long-term contextual information. To imitate real interaction patterns of different speakers, speaker embeddings are also utilized as additional inputs to distinguish the speaker identities during conversational dialogs. To verify the effectiveness of the proposed method, we conduct experiments on the IEMOCAP database. Experimental results demonstrate that our method shows absolute 2.42% performance improvement over the state-of-the-art strategies.
Zheng Lian 0004, Jianhua Tao 0001, Bin Liu 0041, Jian Huang 0014
INTERSPEECH1
2019 Unsupervised Representation Learning with Future Observation Prediction for Speech Emotion Recognition
abstract
Prior works on speech emotion recognition utilize various unsupervised learning approaches to deal with low-resource samples. However, these methods pay less attention to modeling the long-term dynamic dependency, which is important for speech emotion recognition. To deal with this problem, this paper combines the unsupervised representation learning strategy -- Future Observation Prediction (FOP), with transfer learning approaches (such as Fine-tuning and Hypercolumns). To verify the effectiveness of the proposed method, we conduct experiments on the IEMOCAP database. Experimental results demonstrate that our method is superior to currently advanced unsupervised learning strategies.
Zheng Lian 0004, Jianhua Tao 0001, Bin Liu 0041, Jian Huang 0014
INTERSPEECH1
2018 End-to-End Continuous Emotion Recognition from Video Using 3D Convlstm Networks
abstract
Conventional continuous emotion recognition consists of feature extraction step followed by regression step. However, the objective of the two steps is not consistent as they are parted. Besides, there is still no consensus about appropriate emotional features. In this study, we propose an end-to-end continuous emotion recognition framework which merges feature extraction and regressor into a unified system. We employ 3D convolutional networks with Long Short-Term Memory Neutral Network (ConvLSTM) to handle spatiotemporal information for continuous emotion recognition. This model is applied on AVEC 2017 database. The experiment results reveal that ConvLSTM model makes a positive effect on the performance improvement, which outperforms the baseline results for arousal of 0.583 vs 0.525 (baseline) and for valence of 0.h54 vs 0.507.
Jian Huang 0014, Ya Li 0001, Jianhua Tao 0001, Zheng Lian 0004, Jiangyan Yi
ICASSP4