EDBT 2026 Demo / reviewers in the wild / expert
Bin Liu 0041
dblp:35/837-41
· DBLP profile ↗
63ranked-venue papers
7as first author
37since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 39 · 4 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 39 · 6 first-author · 16 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MERBench: A Unified Evaluation Benchmark for Multimodal Emotion RecognitionabstractMultimodal emotion recognition plays a vital role in enhancing user experience in human-computer interaction. Over the past few decades, researchers have developed a range of algorithms and made remarkable progress. While each approach demonstrates certain advantages, inconsistent choices in feature extraction methods, evaluation protocols, and experimental settings have hindered fair comparisons among them. These inconsistencies significantly impede the advancement of the field. To address this issue, we introduce MERBench, a unified evaluation benchmark for multimodal emotion recognition. Our goal is to assess the contributions of several key techniques commonly used in prior studies, such as feature selection, multimodal fusion, robustness analysis, fine-tuning, and pre-training. We believe this work offers clear and comprehensive guidance for future research. Based on the evaluation results of MERBench, we further point out some promising research directions. In addition, we present a new emotion dataset, MER2023, specifically designed for the Chinese language environment. This dataset serves as a benchmark for research in multi-label learning, noise robustness, and semi-supervised learning. Zheng Lian 0004, Licai Sun, Yong Ren 0006, Haiyang Sun 0004, Lan Chen 0005, Bin Liu 0041, Jianhua Tao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | IRNet: Iterative Refinement Network for Noisy Partial Label LearningabstractPartial label learning (PLL) is a typical weakly supervised learning, where each sample is associated with a set of candidate labels. Its basic assumption is that the ground-truth label must be in the candidate set, but this assumption may not be satisfied due to the unprofessional judgment of annotators. Therefore, we relax this assumption and focus on a more general task, noisy PLL, where the ground-truth label may not exist in the candidate set. To address this challenging task, we propose a novel framework called "Iterative Refinement Network (IRNet)", aiming to purify noisy samples through two key modules (i.e., noisy sample detection and label correction). To achieve better performance, we exploit smoothness constraints to reduce prediction errors in these modules. Through theoretical analysis, we prove that IRNet is able to reduce the noise level of the dataset and eventually approximate the Bayes optimal classifier. Meanwhile, IRNet is a plug-in strategy that can be integrated with existing PLL approaches. Experimental results on multiple benchmark datasets show that IRNet outperforms state-of-the-art approaches on noisy PLL. Zheng Lian 0004, Lan Chen 0005, Licai Sun, Bin Liu 0041, Lei Feng 0006, Jianhua Tao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Personality-aware Multimodal Deception Detection with multimodal large language model
Cong Cai, Zhengqi Wen, Xuefei Liu, Jianhua Tao 0001, Bin Liu 0041 |
Pattern Recognit. | 5 |
| 2025 | AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language ModelsabstractThe emergence of multimodal large language models (MLLMs) advances multimodal emotion recognition (MER) to the next level—from naive discriminative tasks to complex emotion understanding with advanced video understanding abilities and natural language description. However, the current community suffers from a lack of large-scale datasets with intensive, descriptive emotion annotations, as well as a multimodal-centric framework to maximize the potential of MLLMs for emotion understanding. To address this, we establish a new benchmark for MLLM-based emotion understanding with a novel dataset (MER-Caption) and a new model (AffectGPT). Utilizing our model-based crowd-sourcing data collection strategy, we construct the largest descriptive emotion dataset to date (by far), featuring over 2K fine-grained emotion categories across 115K samples. We also introduce the AffectGPT model, designed with pre-fusion operations to enhance multimodal integration. Finally, we present MER-UniBench, a unified benchmark with evaluation metrics tailored for typical MER tasks and the free-form, natural language output style of MLLMs. Extensive experimental results show AffectGPT's robust performance across various MER tasks. We have released both the code and the dataset to advance research and development in emotion understanding: https://github.com/zeroQiaoba/AffectGPT. Zheng Lian 0004, Haoyu Chen 0001, Lan Chen 0005, Haiyang Sun 0004, Licai Sun, Yong Ren 0006, Zebang Cheng, Bin Liu 0041, Rui Liu 0008, Xiaojiang Peng, Jiangyan Yi, Jianhua Tao 0001 |
ICML | 8 |
| 2025 | OV-MER: Towards Open-Vocabulary Multimodal Emotion RecognitionabstractMultimodal Emotion Recognition (MER) is a critical research area that seeks to decode human emotions from diverse data modalities. However, existing machine learning methods predominantly rely on predefined emotion taxonomies, which fail to capture the inherent complexity, subtlety, and multi-appraisal nature of human emotional experiences, as demonstrated by studies in psychology and cognitive science. To overcome this limitation, we advocate for introducing the concept of open vocabulary into MER. This paradigm shift aims to enable models to predict emotions beyond a fixed label space, accommodating a flexible set of categories to better reflect the nuanced spectrum of human emotions. To achieve this, we propose a novel paradigm: Open-Vocabulary MER (OV-MER), which enables emotion prediction without being confined to predefined spaces. However, constructing a dataset that encompasses the full range of emotions for OV-MER is practically infeasible; hence, we present a comprehensive solution including a newly curated database, novel evaluation metrics, and a preliminary benchmark. By advancing MER from basic emotions to more nuanced and diverse emotional states, we hope this work can inspire the next generation of MER, enhancing its generalizability and applicability in real-world scenarios. Code and dataset are available at: https://github.com/zeroQiaoba/AffectGPT. Zheng Lian 0004, Haiyang Sun 0004, Licai Sun, Haoyu Chen 0001, Lan Chen 0005, Zhuofan Wen 0001, Hailiang Yao, Bin Liu 0041, Rui Liu 0008, Shan Liang 0007, Ya Li 0001, Jiangyan Yi, Jianhua Tao 0001 |
ICML | 11 |
| 2025 | MER 2025: When Affective Computing Meets Large Language ModelsabstractMER2025 is the third year of our MER series of challenges. Previously, MER2023 (http://merchallenge.cn/mer2023) focused on multi-label learning, noise robustness, and semi-supervised learning, while MER2024 (https://zeroqiaoba.github.io/MER2024-website) introduced a new track dedicated to open-vocabulary emotion recognition. This year, MER2025 centers on the theme ''When Affective Computing Meets Large Language Models (LLMs)''. We aim to shift the paradigm from traditional categorical frameworks reliant on predefined emotion taxonomies to LLM-driven generative methods, offering innovative solutions for more accurate and reliable emotion understanding. The challenge contains four tracks: MER-SEMI focuses on fixed categorical emotion recognition enhanced by semi-supervised learning; MER-FG explores fine-grained emotions, expanding recognition from basic to nuanced emotional states; MER-DES incorporates multimodal cues (beyond emotion words) into predictions to enhance model interpretability; MER-PR reveals whether emotion prediction results can improve personality recognition performance. For the first three tracks, the baseline code is available at MERTools (https://github.com/zeroQiaoba/MERTools) and datasets can be accessed via Hugging Face (https://huggingface.co/datasets/MERChallenge/MER2025). For the last track, the dataset and baseline code are available on GitHub (https://github.com/cai-cong/MER25_personality). Zheng Lian 0004, Rui Liu 0008, Kele Xu, Bin Liu 0041, Xuefei Liu, Yazhou Zhang 0001, Xin Liu 0012, Yong Li 0032, Zebang Cheng, Haolin Zuo, Ziyang Ma 0001, Xiaojiang Peng, Xie Chen 0001, Ya Li 0001, Erik Cambria, Guoying Zhao 0001, Björn W. Schuller, Jianhua Tao 0001 |
ACM Multimedia | 4 |
| 2025 | MDPE: A Multimodal Deception Dataset with Personality and Emotional CharacteristicsabstractDeception detection has garnered increasing attention in recent years due to the significant growth of digital media and heightened ethical and security concerns. It has been extensively studied using multimodal methods, including video, audio, and text. In addition, individual differences in deception production and detection are believed to play a crucial role. Although some studies have utilized individual information such as personality traits to enhance the performance of deception detection, current systems remain limited, partly due to a lack of sufficient datasets for evaluating performance. To address this issue, we introduce a multimodal deception dataset MDPE. Besides deception features, this dataset also includes individual differences information in personality and emotional expression characteristics. It can explore the impact of individual differences on deception behavior. It comprises over 104 hours of deception and emotional videos from 193 subjects. Furthermore, we conducted numerous experiments to provide valuable insights for future deception detection research. MDPE not only supports deception detection, but also provides conditions for tasks such as personality recognition and emotion recognition, and can even study the relationships between them. We believe that MDPE will become a valuable resource for promoting research in the field of affective computing. Cong Cai, Shan Liang 0007, Xuefei Liu, Kang Zhu, Zhengqi Wen, Jianhua Tao 0001, Jizhou Cui, Zhenhua Cheng, Hanzhe Xu, Ruibo Fu, Bin Liu 0041 |
ACM Multimedia | 13 |
| 2025 | Enhancing Multimodal Personality Assessment with LLM-Augmented Hierarchical FusionabstractThis study proposes an LLM-augmented hierarchical fusion framework to enhance multimodal personality and ability assessment for the ACM MULTIMEDIA AVI CHALLENGE 2025, addressing semantic sparsity and cross-modal interaction limitations. We leverage large language models (e.g., Qwen, DeepSeek) to generate psychologically enriched text descriptions, bridging raw transcripts with expert evaluations, and integrate them with audio-visual features through early fusion and multi-path MLP ensembles. Track 1 (personality regression) employs dual-text inputs while Track 2 (multi-label ability prediction) uses parallel regression. Results show significant improvements: 23.1% MSE reduction over text-only baselines in Track 1, and 10.7%/12.5% gains over state-of-the-art fusion in Tracks 1/2, with 31.2% average improvement for cognitive traits (Q3-Q5). The framework demonstrates the effectiveness of semantic enhancement and adaptive fusion, with future work focusing on overfitting mitigation and feature optimization. Longjiang Yang, Zhuofan Wen 0001, Hailiang Yao, Bin Liu 0041, Zheng Lian 0004, Jianhua Tao 0001 |
ACM Multimedia | 9 |
| 2025 | SVFAP: Self-Supervised Video Facial Affect PerceiverabstractVideo-based facial affect analysis has recently attracted increasing attention owing to its critical role in human-computer interaction. Previous studies mainly focus on developing various deep learning architectures and training them in a fully supervised manner. Although significant progress has been achieved by these supervised methods, the longstanding lack of large-scale high-quality labeled data severely hinders their further improvements. Motivated by the recent success of self-supervised learning in computer vision, this paper introduces a self-supervised approach, termed Self-supervised Video Facial Affect Perceiver (SVFAP), to address the dilemma faced by supervised methods. Specifically, SVFAP leverages masked facial video autoencoding to perform self-supervised pre-training on massive unlabeled facial videos. Considering that large spatiotemporal redundancy exists in facial videos, we propose a novel temporal pyramid and spatial bottleneck Transformer as the encoder of SVFAP, which not only largely reduces computational costs but also achieves excellent performance. To verify the effectiveness of our method, we conduct experiments on nine datasets spanning three downstream tasks, including dynamic facial expression recognition, dimensional emotion recognition, and personality recognition. Comprehensive results demonstrate that SVFAP can learn powerful affect-related representations via large-scale self-supervised pre-training and it significantly outperforms previous state-of-the-art methods on all datasets. Licai Sun, Zheng Lian 0004, Haiyang Sun 0004, Bin Liu 0041, Jianhua Tao 0001 |
IEEE Trans. Affect. Comput. | 7 |
| 2025 | Depression Scale Dictionary Decomposition Framework for Multimodal Automatic Depression Level PredictionabstractCurrently, many researchers aim to achieve automatic depression level prediction via speech and video behavior analysis. However, previous works have struggled to decompose audio and video sequences into the information related to and unrelated to depression scores, hindering the model’s perception of depression cues. Besides, previous works implement multimodal fusion using attention mechanisms or linear layers, but failed to simultaneously consider the Euclidean relationship among tokens and the non-Euclidean relationship among channels, which bring limitations in capturing depression cues. In response to the above issues, we propose a depression scale dictionary decomposition framework, which mainly includes a Bidirectional Dictionary Decomposition (BDD) module and a Bidirectional Multimodal Fusion (BMF) module. The BDD module can use the dictionaries generated based on the depression scale to semantically decompose audio and video sequences into the information related to and unrelated to depression scores along token and channel dimensions for promoting depression cue perception. Moreover, considering the respective characteristics of tokens and channels, the BMF module uses linear layers and graph convolution to achieve cross-modal mixing, which is used to aggregate audio and video sequences for predicting depression levels. The validation on AVEC 2013, AVEC 2014 and DAIC-WOZ datasets demonstrates our method’s superiority. Mingyue Niu, Jibing Gong, Bin Liu 0041, Jianhua Tao 0001, Björn W. Schuller |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Pseudo Labels Regularization for Imbalanced Partial-Label LearningabstractPartial-label learning (PLL) is an important branch of weakly supervised learning where the single ground truth resides in a set of candidate labels, while the research rarely considers the label imbalance. A recent study for imbalanced PLL propose that the combinatorial challenge of partial-label learning and long-tail learning lies in matching between a decent marginal prior distribution with drawing the pseudo labels. However, even if the pseudo label matches the prior distribution, the tail classes will still be difficult to learn because the total weight of tail classes is too small. Therefore, we propose a pseudo-label regularization technique specially designed for imbalanced PLL. By punishing the pseudo labels of head classes, our method implements state-of-art under the standardized benchmarks compared to the previous PLL methods. Zheng Lian 0004, Bin Liu 0041, Zerui Chen, Jianhua Tao 0001 |
ICASSP | 3 |
| 2024 | Learning spatial interaction representation with heterogeneous graph convolutional networks for urban land-use inferenceabstractUrban land use is central to urban planning. With the emergence of urban big data and advances in deep learning methods, several studies have leveraged graph convolutional networks (GCNs) with local functional characteristics from points of interest data and spatial features from flow data to infer urban land use. However, these studies cannot distinguish spatial interaction and spatial dependence in terms of conceptualization and modeling mechanisms and overlook the inadequacy of GCNs in modeling spatial interaction. This study proposes a novel framework—a heterogeneous graph convolutional network (HGCN)—to explicitly account for the spatial demand and supply components embedded in spatial interaction data. Several experiments, including 19 different models and datasets from Shenzhen and London, were conducted to validate the proposed framework and its generalizability within the same and different spatial contexts. The HGCN can distinguish heterogeneous mechanisms in supply- and demand-related modalities of spatial interactions, incorporating both spatial interaction and spatial dependence for urban land-use inference. Empowered by HGCN, we found that spatial interaction features play a distinctively crucial role in urban land-use inference compared to local attributes and spatial dependence features. In addition, our findings highlight the superiority of HGCN-based models in boosting performance and enhancing model transferability. Zhaoya Gong, Chenglong Wang 0004, Bin Liu 0041, Zhengzi Zhou |
Int. J. Geogr. Inf. Sci. | 4 |
| 2024 | Efficient Multimodal Transformer With Dual-Level Feature Restoration for Robust Multimodal Sentiment AnalysisabstractWith the proliferation of user-generated online videos, Multimodal Sentiment Analysis (MSA) has attracted increasing attention recently. Despite significant progress, there are still two major challenges on the way towards robust MSA: 1) inefficiency when modeling cross-modal interactions in unaligned multimodal data; and 2) vulnerability to random modality feature missing which typically occurs in realistic settings. In this paper, we propose a generic and unified framework to address them, named Efficient Multimodal Transformer with Dual-Level Feature Restoration (EMT-DLFR). Concretely, EMT employs utterance-level representations from each modality as the global multimodal context to interact with local unimodal features and mutually promote each other. It not only avoids the quadratic scaling cost of previous local-local cross-modal interaction methods but also leads to better performance. To improve model robustness in the incomplete modality setting, on the one hand, DLFR performs low-level feature reconstruction to implicitly encourage the model to learn semantic information from incomplete data. On the other hand, it innovatively regards complete and incomplete data as two different views of one sample and utilizes siamese representation learning to explicitly attract their high-level representations. Comprehensive experiments on three popular datasets demonstrate that our method achieves superior performance in both complete and incomplete modality settings. Licai Sun, Zheng Lian 0004, Bin Liu 0041, Jianhua Tao 0001 |
IEEE Trans. Affect. Comput. | 3 |
| 2024 | Synthesis and Detection Algorithms for Oblique Stripe Noise of Space-Borne Remote Sensing ImagesabstractOblique stripe noise widely appears in remote sensing images after image correction, exhibiting arbitrary tilt angles and parallel distribution. Due to its arbitrary randomness in tilt angles and lengths, oblique stripe noise increases the difficulty of detection compared to vertical or horizontal stripe noise. For the first time, we propose a group of oblique stripe noise synthesis and detection algorithms combining imaging mechanisms and deep learning. To get controllable synthetic oblique stripe noise data for training detection model, two sample augmentation methods are presented by the image correction’s imaging mechanisms with new linear transformation and the generative adversarial network algorithm with Cycle-GAN, respectively. A large-scale simulated stripe noise dataset (SOSD, simulated oblique stripe noise dataset) is simulated using these two methods. A new deep learning detection algorithm (RDOS, Robust detection of oblique stripe Noise) is presented considering the presence of oblique stripe noise. RDOS is trained using both SOSD and a real stripe noise dataset, and it obtains the optimal detection model for testing. The experimental results show that the accuracy reaches 82.93%, the recall rate reaches 85.17%, the F1 score reaches 84.04%, the average precision (AP) reaches 82.34%, and the frames per second (FPS) reaches 33.33. Compared with the general line detection models, our model exceeds ~300% in accuracy and ~60% in speed. In the future, the proposed algorithms have great potential for application in various areas such as quality evaluation, image preprocessing, and engineering problems related to multi-angle linear object augmentation and detection. Binbo Li, Donghai Xie, Yu Wu 0002, Lijuan Zheng, Chongbin Xu, Yibo Fu, Chenglong Wang 0004, Bin Liu 0041 |
IEEE Trans. Geosci. Remote. Sens. | 9 |
| 2024 | PIRNet: Personality-Enhanced Iterative Refinement Network for Emotion Recognition in ConversationabstractEmotion recognition in conversation (ERC) is important for enhancing user experience in human-computer interaction. Unlike vanilla emotion recognition in individual utterances, ERC aims to classify constituent utterances in a dialog into corresponding emotion labels, which makes contextual information crucial. In addition to contextual information, personality traits also affect emotional perception based on psychological findings. Although researchers have proposed several approaches and achieved promising results on ERC, current works in this domain rarely incorporate contextual information and personality influence. To this end, we propose a novel framework to integrate these factors seamlessly, called "Personality-enhanced Iterative Refinement Network (PIRNet)." Specifically, PIRNet is a multistage iterative method. To capture personality influence, PIRNet leverages personality traits to mimic emotional transitions and generates personality-enhanced results. Then we exploit sequence models to capture contextual information in conversations. To verify the effectiveness of our proposed method, we conduct experiments on three benchmark datasets for ERC, that is, IEMOCAP, CMU-MOSI, and CMU-MOSEI. Experimental results demonstrate that our PIRNet succeeds over currently advanced approaches to emotion recognition. Zheng Lian 0004, Bin Liu 0041, Jianhua Tao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | EmotionNAS: Two-stream Neural Architecture Search for Speech Emotion Recognition
Haiyang Sun 0004, Zheng Lian 0004, Bin Liu 0041, Jianhua Tao 0001, Licai Sun, Cong Cai, Meng Wang 0001 |
INTERSPEECH | 3 |
| 2023 | MER 2023: Multi-label Learning, Modality Robustness, and Semi-Supervised LearningabstractThe first Multimodal Emotion Recognition Challenge (MER 2023)1 was successfully held at ACM Multimedia. The challenge focuses on system robustness and consists of three distinct tracks: (1) MER-MULTI, where participants are required to recognize both discrete and dimensional emotions; (2) MER-NOISE, in which noise is added to test videos for modality robustness evaluation; (3) MER-SEMI, which provides a large amount of unlabeled samples for semi-supervised learning. In this paper, we introduce the motivation behind this challenge, describe the benchmark dataset, and provide some statistics about participants. To continue using this dataset after MER 2023, please sign a new End User License Agreement2 and send it to our official email address3. We believe this high-quality dataset can become a new benchmark in multimodal emotion recognition, especially for the Chinese research community. Zheng Lian 0004, Haiyang Sun 0004, Licai Sun, Jinming Zhao, Ye Liu 0010, Bin Liu 0041, Jiangyan Yi, Meng Wang 0001, Erik Cambria, Guoying Zhao 0001, Björn W. Schuller, Jianhua Tao 0001 |
ACM Multimedia | 12 |
| 2023 | MAE-DFER: Efficient Masked Autoencoder for Self-supervised Dynamic Facial Expression RecognitionabstractDynamic facial expression recognition (DFER) is essential to the development of intelligent and empathetic machines. Prior efforts in this field mainly fall into supervised learning paradigm, which is severely restricted by the limited labeled data in existing datasets. Inspired by recent unprecedented success of masked autoencoders (e.g., VideoMAE), this paper proposes MAE-DFER, a novel self-supervised method which leverages large-scale self-supervised pre-training on abundant unlabeled data to largely advance the development of DFER. Since the vanilla Vision Transformer (ViT) employed in VideoMAE requires substantial computation during fine-tuning, MAE-DFER develops an efficient local-global interaction Transformer (LGI-Former) as the encoder. Moreover, in addition to the standalone appearance content reconstruction in VideoMAE, MAE-DFER also introduces explicit temporal facial motion modeling to encourage LGI-Former to excavate both static appearance and dynamic motion information. Extensive experiments on six datasets show that MAE-DFER consistently outperforms state-of-the-art supervised methods by significant margins (e.g., +6.30% UAR on DFEW and +8.34% UAR on MAFW), verifying that it can learn powerful dynamic facial representations via large-scale self-supervised pre-training. Besides, it has comparable or even better performance than VideoMAE, while largely reducing the computational cost (about 38% FLOPs). We believe MAE-DFER has paved a new way for the advancement of DFER and can inspire more relevant research in this field and even other related tasks. Codes and models are publicly available at https://github.com/sunlicai/MAE-DFER. Licai Sun, Zheng Lian 0004, Bin Liu 0041, Jianhua Tao 0001 |
ACM Multimedia | 3 |
| 2023 | Integrating VideoMAE based model and Optical Flow for Micro- and Macro-expression SpottingabstractThe task of interval localization of macro- and micro-expression in long videos has a wide range of applications in the field of human-computer interaction. Compared with macro-expression, micro-expression has shorter duration, lower intensity, and smaller number of samples, which make them more difficult to spot accurately in long videos. In this paper, we propose a pre-trained model combined with the optical flow method to improve the accuracy and robustness of macro- and micro-expression spotting. Firstly, self-supervised pre-training is performed on rich unlabeled data based on VideoMAE. Then, multiple models are trained on the datasets SAMM-LV and CAS(ME)³ for macro- and micro-expression with different fine-grains. Finally, different lengths of slices are generated based on the models with different fine-grains, and the optimal matching method through the combination of model fine-grainedness and slice lengths is explored. At the same time, macro- and micro-expression generating regions were spotted using the optical flow method, fused with the model outputs to supplement the spatio-temporal information not captured by the model and to exclude the interference of non-interested regions. We evaluated the performance of our method on the MEGC2023 testset (consisting of 10 long videos from SAMM and 20 long videos from CAS(ME)3) and won first place in the MEGC2023 Challenge. The results demonstrate the effectiveness of the method. Licai Sun, Zheng Lian 0004, Bin Liu 0041, Haiyang Sun 0004, Jianhua Tao 0001 |
ACM Multimedia | 5 |
| 2023 | ALIM: Adjusting Label Importance Mechanism for Noisy Partial Label LearningabstractNoisy partial label learning (noisy PLL) is an important branch of weakly supervised learning. Unlike PLL where the ground-truth label must conceal in the candidate label set, noisy PLL relaxes this constraint and allows the ground-truth label may not be in the candidate label set. To address this challenging problem, most of the existing works attempt to detect noisy samples and estimate the ground-truth label for each noisy sample. However, detection errors are unavoidable. These errors can accumulate during training and continuously affect model optimization. To this end, we propose a novel framework for noisy PLL with theoretical interpretations, called ``Adjusting Label Importance Mechanism (ALIM)''. It aims to reduce the negative impact of detection errors by trading off the initial candidate set and model outputs. ALIM is a plug-in strategy that can be integrated with existing PLL approaches. Experimental results on multiple benchmark datasets demonstrate that our method can achieve state-of-the-art performance on noisy PLL. Our code is available at: https://github.com/zeroQiaoba/ALIM. Zheng Lian 0004, Lei Feng 0006, Bin Liu 0041, Jianhua Tao 0001 |
NeurIPS | 4 |
| 2023 | VRA: Variational Rectified Activation for Out-of-distribution DetectionabstractOut-of-distribution (OOD) detection is critical to building reliable machine learning systems in the open world. Researchers have proposed various strategies to reduce model overconfidence on OOD data. Among them, ReAct is a typical and effective technique to deal with model overconfidence, which truncates high activations to increase the gap between in-distribution and OOD. Despite its promising results, is this technique the best choice? To answer this question, we leverage the variational method to find the optimal operation and verify the necessity of suppressing abnormally low and high activations and amplifying intermediate activations in OOD detection, rather than focusing only on high activations like ReAct. This motivates us to propose a novel technique called ``Variational Rectified Activation (VRA)'', which simulates these suppression and amplification operations using piecewise functions. Experimental results on multiple benchmark datasets demonstrate that our method outperforms existing post-hoc strategies. Meanwhile, VRA is compatible with different scoring functions and network architectures. Our code is available at https://github.com/zeroQiaoba/VRA. Zheng Lian 0004, Bin Liu 0041, Jianhua Tao 0001 |
NeurIPS | 3 |
| 2023 | GCNet: Graph Completion Network for Incomplete Multimodal Learning in ConversationabstractConversations have become a critical data format on social media platforms. Understanding conversation from emotion, content and other aspects also attracts increasing attention from researchers due to its widespread application in human-computer interaction. In real-world environments, we often encounter the problem of incomplete modalities, which has become a core issue of conversation understanding. To address this problem, researchers propose various methods. However, existing approaches are mainly designed for individual utterances rather than conversational data, which cannot fully exploit temporal and speaker information in conversations. To this end, we propose a novel framework for incomplete multimodal learning in conversations, called "Graph Complete Network (GCNet)," filling the gap of existing works. Our GCNet contains two well-designed graph neural network-based modules, "Speaker GNN" and "Temporal GNN," to capture temporal and speaker dependencies. To make full use of complete and incomplete data, we jointly optimize classification and reconstruction tasks in an end-to-end manner. To verify the effectiveness of our method, we conduct experiments on three benchmark conversational datasets. Experimental results demonstrate that our GCNet is superior to existing state-of-the-art approaches in incomplete multimodal learning. Zheng Lian 0004, Lan Chen 0005, Licai Sun, Bin Liu 0041, Jianhua Tao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | SMIN: Semi-Supervised Multi-Modal Interaction Network for Conversational Emotion RecognitionabstractConversational emotion recognition is a crucial research topic in human-computer interactions. Due to the heavy annotation cost and inevitable label ambiguity, collecting large amounts of labeled data is challenging and expensive, which restricts the performance of current fully-supervised methods in this domain. To address this problem, researchers attempt to distill knowledge from unlabeled data via semi-supervised learning. However, most of these semi-supervised methods ignore multimodal interactive information, although recent works have proven that such interactive information is essential for emotion recognition. To this end, we propose a novel framework to seamlessly integrate semi-supervised learning with multimodal interactions, called “Semi-supervised Multi-modal Interaction Network (SMIN)”. SMIN contains two well-designed semi-supervised modules, “Intra-modal Interactive Module (IIM)” and “Cross-modal Interactive Module (CIM)” to learn intra- and cross-modal interactions. These two modules leverage additional unlabeled data to extract emotion-salient representations. To capture additional contextual information, we utilize the hierarchical recurrent networks followed with the hybrid fusion strategy to integrate multimodal features. These multimodal features are further utilized for conversational emotion recognition. Experimental results on four benchmark datasets (i.e., IEMOCAP, MELD, CMU-MOSI and CMU-MOSEI) demonstrate that SMIN succeeds over existing state-of-the-art strategies on emotion recognition. Zheng Lian 0004, Bin Liu 0041, Jianhua Tao 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | Multimodal Spatiotemporal Representation for Automatic Depression Level DetectionabstractPhysiological studies have shown that there are some differences in speech and facial activities between depressive and healthy individuals. Based on this fact, we propose a novel spatio-temporal attention (STA) network and a multimodal attention feature fusion (MAFF) strategy to obtain the multimodal representation of depression cues for predicting the individual depression level. Specifically, we first divide the speech amplitude spectrum/video into fixed-length segments and input these segments into the STA network, which not only integrates the spatial and temporal information through attention mechanism, but also emphasizes the audio/video frames related to depression detection. The audio/video segment-level feature is obtained from the output of the last full connection layer of the STA network. Second, this article employs the eigen evolution pooling method to summarize the changes of each dimension of the audio/video segment-level features to aggregate them into the audio/video level feature. Third, the multimodal representation with modal complementary information is generated using the MAFF and inputs into the support vector regression predictor for estimating depression severity. Experimental results on the AVEC2013 and AVEC2014 depression databases illustrate the effectiveness of our method. Mingyue Niu, Jianhua Tao 0001, Bin Liu 0041, Jian Huang 0014, Zheng Lian 0004 |
IEEE Trans. Affect. Comput. | 3 |
| 2023 | Dense Modality Interaction Network for Audio-Visual Event LocalizationabstractHuman perception systems can integrate audio and visual information automatically to obtain a profound understanding of real-world events. Accordingly, fusing audio and visual contents is important to solve the audio-visual event (AVE) localization problem. Although most existing works have fused audio and visual modalities to explore their relationship with attention-based networks, we can delve into their relationship more deeply to improve the fusion capability of the two modalities. In this paper, we propose a dense modality interaction network (DMIN) to elegantly leverage audio and visual information by integrating two novel modules, namely, the audio-guided triplet attention (AGTA) module and the dense inter-modality attention (DIMA) module. The AGTA module enables audio information to guide the network to pay more attention to event-relevant visual regions. This guidance is conducted in the channel, temporal, and spatial dimensions, which emphasize informative features, temporal relationships and spatial regions, to boost the capacity of representations. Furthermore, the DIMA module establishes the dense-relationship between audio and visual modalities. Specifically, the DIMA module leverages the information of all channel pairs of audio and visual features to formulate the cross-modality attention weight, which is superior to the multi-head attention module that uses limited information. Moreover, a novel unimodal discrimination loss (UDL) is introduced to exploit the unimodal and fused features together for more exact AVE localization. The experimental results show that our method is remarkably superior to the state-of-the-art methods in fully- and weakly-supervised AVE settings. To further evaluate the model's ability to build audio-visual connections, we design a dense cross modality relation network (DCMR) to solve the cross-modality localization task. DCMR is a simple deformation of a DMIN, and the experimental results further illustrate that DIMA can explore denser relationships between the two modalities. Code is available at https://github.com/weizequan/DMIN.git. Weize Quan, Bin Liu 0041, Dong-Ming Yan 0001 |
IEEE Trans. Multim. | 5 |
| 2022 | End-to-End Network Based on Transformer for Automatic Detection of Covid-19abstractThe novel coronavirus disease (COVID-19) was declared a pandemic by the World Health Organization. The cumulative number of deaths is more than 4.8 million. Epidemiology experts concur that mass testing is essential for isolating infected individuals, contact tracing, and slowing the progression of the virus. In recent months, some machine learning methods have been proposed utilizing audio cues for COVID-19 detection. However, many works are based on hand-crafted features and deep features to detect COVID-19. There is no evidence that these features are optimal for COVID-19 detection. Therefore, we proposed an end-to-end network based on transformer for automatic detection of COVID-19. It directly learns features from the raw waveform for end-to-end learning, rather than extracting features in advance. We propose a feature extraction module to automatically extract features. And we use the transformer architectures to model the dependencies between the extracted features. It is the first end-to-end learning based on raw waveform for COVID-19 detection. Experiments on COUGHVID dataset show that our method has achieved competitive results. Cong Cai, Bin Liu 0041, Jianhua Tao 0001, Zhengkun Tian |
ICASSP | 2 |
| 2022 | Depressioner: Facial dynamic representation for automatic depression level prediction
Mingyue Niu, Ya Li 0001, Bin Liu 0041 |
Expert Syst. Appl. | 4 |
| 2021 | Multi-Scale and Multi-Region Facial Discriminative Representation for Automatic Depression Level PredictionabstractPhysiological studies have shown that differences in facial activities between depressed patients and normal individuals are manifested in different local facial regions and the durations of these activities are not the same. But most previous works extract features from the entire facial region at a fixed time scale to predict the individual depression level. Thus, they are inadequate in capturing dynamic facial changes. For these reasons, we propose a multi-scale and multi-region fa-cial dynamic representation method to improve the prediction performance. In particular, we firstly use multiple time scales to divide the original long-term video into segments containing different facial regions. Secondly, the segment-level feature is extracted by 3D convolution neural network to characterize the facial activities with different durations in different facial regions. Thirdly, this paper adopts eigen evolution pooling and gradient boosting decision tree to aggregate these segment-level features and select discriminative elements to generate the video-level feature. Finally, the depression level is predicted using support vector regression. Experiments are conducted on AVEC2013 and AVEC2014. The results demonstrate that our method achieves better performance than the previous works. Mingyue Niu, Jianhua Tao 0001, Bin Liu 0041 |
ICASSP | 3 |
| 2021 | Multimodal Cross- and Self-Attention Network for Speech Emotion RecognitionabstractSpeech Emotion Recognition (SER) requires a thorough understanding of both the linguistic content of an utterance (i.e., textual information) and how the speaker utters it (i.e., acoustic information). The one vital challenge in SER is how to effectively fuse these two kinds of information. In this paper, we propose a novel Multimodal Cross- and Self-Attention Network (MCSAN) to tackle this problem. The core of MCSAN is to employ the parallel cross- and self-attention modules to explicitly model both inter- and intra-modal interactions of audio and text. Specifically, the cross-attention module utilizes the cross-attention mechanism to guide one modality to attend to the other modality and update the features accordingly. Similarly, the self-attention module employs the self-attention mechanism to propagate information within each modality. We evaluate MCSAN on two benchmark datasets, IEMOCAP and MELD. Experimental results demonstrate that our proposed model achieves state-of-the-art performance on both datasets. Licai Sun, Bin Liu 0041, Jianhua Tao 0001, Zheng Lian 0004 |
ICASSP | 2 |
| 2021 | TDCA-Net: Time-Domain Channel Attention Network for Depression Detection
Cong Cai, Mingyue Niu, Bin Liu 0041, Jianhua Tao 0001, Xuefei Liu |
Interspeech | 3 |
| 2021 | DECN: Dialogical emotion correction network for conversational emotion recognition
Zheng Lian 0004, Bin Liu 0041, Jianhua Tao 0001 |
Neurocomputing | 2 |
| 2021 | A time-frequency channel attention and vectorization network for automatic depression level prediction
Mingyue Niu, Bin Liu 0041, Jianhua Tao 0001 |
Neurocomputing | 2 |
| 2021 | Gated Recurrent Fusion With Joint Training Framework for Robust End-to-End Speech RecognitionabstractThe joint training framework for speech enhancement and recognition methods have obtained quite good performances for robust end-to-end automatic speech recognition (ASR). However, these methods only utilize the enhanced feature as the input of the speech recognition component, which are affected by the speech distortion problem. In order to address this problem, this paper proposes a gated recurrent fusion (GRF) method with joint training framework for robust end-to-end ASR. The GRF algorithm is used to dynamically combine the noisy and enhanced features. Therefore, the GRF can not only remove the noise signals from the enhanced features, but also learn the raw fine structures from the noisy features so that it can alleviate the speech distortion. The proposed method consists of speech enhancement, GRF and speech recognition. Firstly, the mask based speech enhancement network is applied to enhance the input speech. Secondly, the GRF is applied to address the speech distortion problem. Thirdly, to improve the performance of ASR, the state-of-the-art speech transformer algorithm is used as the speech recognition component. Finally, the joint training framework is utilized to optimize these three components, simultaneously. Our experiments are conducted on an open-source Mandarin speech corpus called AISHELL-1. Experimental results show that the proposed method achieves the relative character error rate (CER) reduction of 10.04% over the conventional joint enhancement and transformer method only using the enhanced features. Especially for the low signal-to-noise ratio (0 dB), our proposed method can achieves better performances with 12.67% CER reduction, which suggests the potential of our proposed method. Cunhang Fan, Jiangyan Yi, Jianhua Tao 0001, Zhengkun Tian, Bin Liu 0041, Zhengqi Wen |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2021 | $F_0$-Noise-Robust Glottal Source and Vocal Tract Analysis Based on ARX-LF ModelabstractThis paper proposes a robust automatic speech analysis method based on a source-filter model constructed of an Auto-Regressive eXogenous (ARX) model and the Liljencrants-Fant (LF) model. The proposed method estimates glottal source waveform and vocal tract shape parameters using an analysis-by-synthesis approach. Structurally, the first step is to initialize the glottal source parameters using the inverse filter method, and the second step is to simultaneously estimate the glottal source waveform and the vocal tract shape parameters using an analysis-by-synthesis approach with an iterative algorithm. The proposed method was verified on synthetic voices with different glottal noise (signal to noise ratio) from 0 dB to 50 dB and different fundamental frequency ( F_0 ) from 80 Hz to 320 Hz levels. The results show that the proposed method achieved a much higher estimation accuracy than that of the state-of-the-art inverse filtering methods on both different glottal noise and different F_0 levels. Jianhua Tao 0001, Donna Erickson, Bin Liu 0041, Masato Akagi |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | CTNet: Conversational Transformer Network for Emotion RecognitionabstractEmotion recognition in conversation is a crucial topic for its widespread applications in the field of human-computer interactions. Unlike vanilla emotion recognition of individual utterances, conversational emotion recognition requires modeling both context-sensitive and speaker-sensitive dependencies. Despite the promising results of recent works, they generally do not leverage advanced fusion techniques to generate the multimodal representations of an utterance. In this way, they have limitations in modeling the intra-modal and cross-modal interactions. In order to address these problems, we propose a multimodal learning framework for conversational emotion recognition, called conversational transformer network (CTNet). Specifically, we propose to use the transformer-based structure to model intra-modal and cross-modal interactions among multimodal features. Meanwhile, we utilize word-level lexical features and segment-level acoustic features as the inputs, thus enabling us to capture temporal information in the utterance. Additionally, to model context-sensitive and speaker-sensitive dependencies, we propose to use the multihead attention based bi-directional GRU component and speaker embeddings. Experimental results on the IEMOCAP and MELD datasets demonstrate the effectiveness of the proposed method. Our method shows an absolute 2.1~6.2% performance improvement on weighted average F1 over state-of-the-art strategies. Zheng Lian 0004, Bin Liu 0041, Jianhua Tao 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Learning long-term temporal contexts using skip RNN for continuous emotion recognitionabstractOne of the most critical issues in human-computer interaction applications is recognizing human emotions based on speech. In recent years, the challenging problem of cross-corpus speech emotion recognition (SER) has generated extensive research. Nevertheless, the domain discrepancy between training data and testing data remains a major challenge to achieving improved system performance. This paper introduces a novel multi-scale discrepancy adversarial (MSDA) network for conducting multiple timescales domain adaptation for cross-corpus SER, i. e., integrating domain discriminators of hierarchical levels into the emotion recognition framework to mitigate the gap between the source and target domains. Specifically, we extract two kinds of speech features, i.e., handcraft features and deep features, from three timescales of global, local, and hybrid levels. In each timescale, the domain discriminator and the emotion classifier compete against each other to learn features that minimize the discrepancy between the two domains by fooling the discriminator. Extensive experiments on cross-corpus and cross-language SER were conducted on a combination dataset that combines one Chinese dataset and two English datasets commonly used in SER. The MSDA is affected by the strong discriminate power provided by the adversarial process, where three discriminators are working in tandem with an emotion classifier. Accordingly, the MSDA achieves the best performance over all other baseline methods. The proposed architecture was tested on a combination of one Chinese and two English datasets. The experimental results demonstrate the superiority of our powerful discriminative model for solving cross-corpus SER. Jian Huang 0014, Bin Liu 0041, Jianhua Tao 0001 |
Virtual Real. Intell. Hardw. | 2 |
| 2021 | Review of micro-expression spotting and recognition in video sequencesabstractFacial micro-expressions are short and imperceptible expressions that involuntarily reveal the true emotions that a person may be attempting to suppress, hide, disguise, or conceal. Such expressions can reflect a person's real emotions and have a wide range of application in public safety and clinical diagnosis. The analysis of facial micro-expressions in video sequences through computer vision is still relatively recent. In this research, a comprehensive review on the topic of spotting and recognition used in microexpression analysis databases and methods, is conducted, and advanced technologies in this area are summarized. In addition, we discuss challenges that remain unresolved alongside future work to be completed in the field of micro-expression analysis. Hang Pan 0001, Lun Xie, Bin Liu 0041, Jianhua Tao 0001 |
Virtual Real. Intell. Hardw. | 4 |
| 2020 | Multimodal Transformer Fusion for Continuous Emotion RecognitionabstractMultimodal fusion increases the performance of emotion recognition because of the complementarity of different modalities. Compared with decision level and feature level fusion, model level fusion makes better use of the advantages of deep neural networks. In this work, we utilize the Transformer model to fuse audio-visual modalities on the model level. Specifically, the multi-head attention produces multimodal emotional intermediate representations from common semantic feature space after encoding audio and visual modalities. Meanwhile, it also can learn long-term temporal dependencies with self-attention mechanism effectively. The experiments, on the AVEC 2017 database, shows the superiority of model level fusion than other fusion strategies. Moreover, we combine the Transformer model and LSTM to further improve the performance, which achieves better results than other methods. Jian Huang 0014, Jianhua Tao 0001, Bin Liu 0041, Zheng Lian 0004, Mingyue Niu |
ICASSP | 3 |
| 2020 | Gated Recurrent Fusion of Spatial and Spectral Features for Multi-Channel Speech Separation with Deep Embedding Representations
Cunhang Fan, Jianhua Tao 0001, Bin Liu 0041, Jiangyan Yi, Zhengqi Wen |
INTERSPEECH | 3 |
| 2020 | Joint Training for Simultaneous Speech Denoising and Dereverberation with Deep Embedding Representations
Cunhang Fan, Jianhua Tao 0001, Bin Liu 0041, Jiangyan Yi, Zhengqi Wen |
INTERSPEECH | 3 |
| 2020 | Learning Utterance-Level Representations with Label Smoothing for Speech Emotion Recognition
Jian Huang 0014, Jianhua Tao 0001, Bin Liu 0041, Zheng Lian 0004 |
INTERSPEECH | 3 |
| 2020 | Comparison of Glottal Source Parameter Values in Emotional VowelsabstractSince glottal source plays an important role for expressing emotions in speech, it is crucial to compare a set of glottal source parameter values to find differences in these expressions of emotions for emotional speech recognition and synthesis. This paper focuses on comparing a set of glottal source parameter values among varieties of emotional vowels /a/ (joy, neutral, anger, and sadness) using an improved ARX-LF model algorithm. The set of glottal source parameters included in the comparison were T_p, T_e, T_a, E_e, and F_0(1/T_0) in the LF model; parameter values were divided into 5 levels according to that of neutral vowel. Results showed that each emotion has its own levels for each set of the glottal source parameter value. These findings could be used for emotional speech recognition and synthesis. Jianhua Tao 0001, Bin Liu 0041, Donna Erickson, Masato Akagi |
INTERSPEECH | 3 |
| 2020 | Context-Dependent Domain Adversarial Neural Network for Multimodal Emotion Recognition
Zheng Lian 0004, Jianhua Tao 0001, Bin Liu 0041, Jian Huang 0014, Zhanlei Yang, Rongjun Li |
INTERSPEECH | 3 |
| 2020 | Conversational Emotion Recognition Using Self-Attention Mechanisms and Graph Neural Networks
Zheng Lian 0004, Jianhua Tao 0001, Bin Liu 0041, Jian Huang 0014, Zhanlei Yang, Rongjun Li |
INTERSPEECH | 3 |
| 2020 | Hybrid Network Feature Extraction for Depression Assessment from SpeechabstractA fast-growing area of mental health research is the search for speech-based objective markers for conditions such as depression.One vital challenge in the development of speech-based depression severity assessment systems is the extraction of depression-relevant features from speech signals.In order to deliver more comprehensive feature representation, we herein explore the benefits of a hybrid network that encodes depressionrelated characteristics in speech for the task of depression severity assessment.The proposed network leverages self-attention networks (SAN) trained on low-level acoustic features and deep convolutional neural networks (DCNN) trained on 3D Log-Mel spectrograms.The feature representations learnt in the SAN and DCNN are concatenated and average pooling is exploited to aggregate complementary segment-level features.Finally, support vector regression is applied to predict a speaker's Beck Depression Inventory-II score.Experiments based on a subset of the Audio-Visual Depressive Language Corpus, as used in the 2013 and 2014 Audio/Visual Emotion Challenges, demonstrate the effectiveness of our proposed hybrid approach. Ziping Zhao 0001, Nicholas Cummins, Bin Liu 0041, Haishuai Wang, Jianhua Tao 0001, Björn W. Schuller |
INTERSPEECH | 4 |
| 2020 | End-to-End Post-Filter for Speech Separation With Deep Attention Fusion FeaturesabstractIn this article, we propose an end-to-end post-filter method with deep attention fusion features for monaural speaker-independent speech separation. At first, a time-frequency domain speech separation method is applied as the pre-separation stage. The aim of pre-separation stage is to separate the mixture preliminarily. Although this stage can separate the mixture, it still contains the residual interference. In order to enhance the pre-separated speech and improve the separation performance further, the end-to-end post-filter (E2EPF) with deep attention fusion features is proposed. The E2EPF can make full use of the prior knowledge of the pre-separated speech, which contributes to speech separation. It is a fully convolutional speech separation network and uses the waveform as the input features. Firstly, the 1-D convolutional layer is utilized to extract the deep representation features for the mixture and pre-separated signals in the time domain. Secondly, to pay more attention to the outputs of the pre-separation stage, an attention module is applied to acquire deep attention fusion features, which are extracted by computing the similarity between the mixture and the pre-separated speech. These deep attention fusion features are conducive to reduce the interference and enhance the pre-separated speech. Finally, these features are sent to the post-filter to estimate each target signals. Experimental results on the WSJ0-2mix dataset show that the proposed method outperforms the state-of-the-art speech separation method. Compared with the pre-separation method, our proposed method can acquire 64.1%, 60.2%, 25.6% and 7.5% relative improvements in scale-invariant source-to-noise ratio (SI-SNR), the signal-to-distortion ratio (SDR), the perceptual evaluation of speech quality (PESQ) and the short-time objective intelligibility (STOI) measures, respectively. Cunhang Fan, Jianhua Tao 0001, Bin Liu 0041, Jiangyan Yi, Zhengqi Wen, Xuefei Liu |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | Efficient Modeling of Long Temporal Contexts for Continuous Emotion RecognitionabstractContinuous emotion recognition is a challenging task due to its difficulty in modeling long-term contexts dependencies. Prior researches have exploited emotional temporal contexts from two perspectives, which are based on feature representations and emotional models. In this paper, we explore the model based approaches for continuous emotion recognition. Specifically, three temporal models including LSTM, TDNN and multi-head attention models are utilized to learn long-term contexts dependencies based on short-term feature representations. The temporal information learned by the temporal models allows the network to more easily exploit the slow changing dynamics between emotional states. Our experimental results demonstrate that the temporal models can model emotional long-term dynamic information effectively. Multi-head attention model achieves best performance among three models and multi-model combination models further improve the performance of continuous emotion recognition significantly. Jian Huang 0014, Jianhua Tao 0001, Bin Liu 0041, Zhen Lian, Mingyue Niu |
ACII | 3 |
| 2019 | Loss and Double-edge-triggered Detector for Robust Small-footprint Keyword SpottingabstractKeyword spotting (KWS) system constitutes a critical component of human-computer interfaces, which detects the specific keyword from a continuous stream of audio. The goal of KWS is providing a high detection accuracy at a low false alarm rate while having small memory and computation requirements. The DNN-based KWS system faces a large class imbalance during training because the amount of data available for the keyword is usually much less than the background speech, which overwhelms training and leads to a degenerate model. In this paper, we explore the focal loss for the training of a small-footprint KWS system. It can automatically down-weight the contribution of easy samples during training and focus the model on hard samples, which naturally solves the class imbalance and allows us to efficiently utilize all data available. Furthermore, many keywords of Chinese conversational assistants are repeated words due to the idiomatic usage, such as `XIAO DU XIAO DU'. We propose a double-edge-triggered detecting method for the repeated keyword, which significantly reduces the false alarm rate relative to the single threshold method. Systematic experiments demonstrate significant further improvements compared to the baseline system. Bin Liu 0041, Shuai Nie 0001, Shan Liang 0007, Zhanlei Yang |
ICASSP | 1 |
| 2019 | Discriminative Learning for Monaural Speech Separation Using Deep Embedding FeaturesabstractDeep clustering (DC) and utterance-level permutation invariant training (uPIT) have been demonstrated promising for speakerindependent speech separation.DC is usually formulated as two-step processes: embedding learning and embedding clustering, which results in complex separation pipelines and a huge obstacle in directly optimizing the actual separation objectives.As for uPIT, it only minimizes the chosen permutation with the lowest mean square error, doesn't discriminate it with other permutations.In this paper, we propose a discriminative learning method for speaker-independent speech separation using deep embedding features.Firstly, a DC network is trained to extract deep embedding features, which contain each source's information and have an advantage in discriminating each target speakers.Then these features are used as the input for uPIT to directly separate the different sources.Finally, uPIT and DC are jointly trained, which directly optimizes the actual separation objectives.Moreover, in order to maximize the distance of each permutation, the discriminative learning is applied to fine tuning the whole model.Our experiments are conducted on WSJ0-2mix dataset.Experimental results show that the proposed models achieve better performances than DC and uPIT for speaker-independent speech separation. Cunhang Fan, Bin Liu 0041, Jianhua Tao 0001, Jiangyan Yi, Zhengqi Wen |
INTERSPEECH | 2 |
| 2019 | Conversational Emotion Analysis via Attention MechanismsabstractDifferent from the emotion recognition in individual utterances, we propose a multimodal learning framework using relation and dependencies among the utterances for conversational emotion analysis. The attention mechanism is applied to the fusion of the acoustic and lexical features. Then these fusion representations are fed into the self-attention based bi-directional gated recurrent unit (GRU) layer to capture long-term contextual information. To imitate real interaction patterns of different speakers, speaker embeddings are also utilized as additional inputs to distinguish the speaker identities during conversational dialogs. To verify the effectiveness of the proposed method, we conduct experiments on the IEMOCAP database. Experimental results demonstrate that our method shows absolute 2.42% performance improvement over the state-of-the-art strategies. Zheng Lian 0004, Jianhua Tao 0001, Bin Liu 0041, Jian Huang 0014 |
INTERSPEECH | 3 |
| 2019 | Unsupervised Representation Learning with Future Observation Prediction for Speech Emotion RecognitionabstractPrior works on speech emotion recognition utilize various unsupervised learning approaches to deal with low-resource samples. However, these methods pay less attention to modeling the long-term dynamic dependency, which is important for speech emotion recognition. To deal with this problem, this paper combines the unsupervised representation learning strategy -- Future Observation Prediction (FOP), with transfer learning approaches (such as Fine-tuning and Hypercolumns). To verify the effectiveness of the proposed method, we conduct experiments on the IEMOCAP database. Experimental results demonstrate that our method is superior to currently advanced unsupervised learning strategies. Zheng Lian 0004, Jianhua Tao 0001, Bin Liu 0041, Jian Huang 0014 |
INTERSPEECH | 3 |
| 2019 | Jointly Adversarial Enhancement Training for Robust End-to-End Speech Recognition
Bin Liu 0041, Shuai Nie 0001, Shan Liang 0007, Meng Yu 0003, Lianwu Chen, Shouye Peng, Changliang Li |
INTERSPEECH | 1 |
| 2019 | Automatic Depression Level Detection via ℓp-Norm Pooling
Mingyue Niu, Jianhua Tao 0001, Bin Liu 0041, Cunhang Fan |
INTERSPEECH | 3 |
| 2018 | Boosting Noise Robustness of Acoustic Model via Deep Adversarial TrainingabstractIn realistic environments, speech is usually interfered by various noise and reverberation, which dramatically degrades the performance of automatic speech recognition (ASR) systems. To alleviate this issue, the commonest way is to use a well-designed speech enhancement approach as the front-end of ASR. However, more complex pipelines, more computations and even higher hardware costs (microphone array) are additionally consumed for this kind of methods. In addition, speech enhancement would result in speech distortions and mismatches to training. In this paper, we propose an adversarial training method to directly boost noise robustness of acoustic model. Specifically, a jointly compositional scheme of generative adversarial net (GAN) and neural network-based acoustic model (AM) is used in the training phase. GAN is used to generate clean feature representations from noisy features by the guidance of a discriminator that tries to distinguish between the true clean signals and generated signals. The joint optimization of generator, discriminator and AM concentrates the strengths of both GAN and AM for speech recognition. Systematic experiments on CHiME-4 show that the proposed method significantly improves the noise robustness of AM and achieves the average relative error rate reduction of 23.38% and 11.54% on the development and test set, respectively. Bin Liu 0041, Shuai Nie 0001, Dengfeng Ke, Shan Liang 0007 |
ICASSP | 1 |
| 2018 | Stochastic Multiple Choice Learning for Acoustic ModelingabstractEven for deep neural networks, it is still a challenging task to indiscriminately model thousands of fine-grained senones only by one model. Ensemble learning is a well-known technique that is capable of concentrating the strengths of different models to facilitate the complex task. In addition, the phones may be spontaneously aggregated into several clusters due to the intuitive perceptual properties of speech, such as vowels and consonants. However, a typical ensemble learning scheme usually trains each submodular independently and doesn't explicitly consider the internal relation of data, which is hardly expected to improve the classification performance of fine-grained senones. In this paper, we use a novel training schedule for DNN-based ensemble acoustic model. In the proposed training schedule, all submodels are jointly trained to cooperatively optimize the loss objective by a Stochastic Multiple Choice Learning approach. It results in that different submodels have specialty capacities for modeling senones with different properties. Systematic experiments show that the proposed model is competitive with the dominant DNN-based acoustic models in the TIMIT and THCHS-30 recognition tasks. Bin Liu 0041, Shuai Nie 0001, Shan Liang 0007, Zhanlei Yang |
IJCNN | 1 |
| 2018 | Deep Noise Tracking Network: A Hybrid Signal Processing/Deep Learning Approach to Speech Enhancement
Shuai Nie 0001, Shan Liang 0007, Bin Liu 0041, Jianhua Tao 0001 |
INTERSPEECH | 3 |
| 2017 | A novel pitch extraction based on jointly trained deep BLSTM Recurrent Neural Networks with bottleneck featuresabstractPitch is an important characteristic of speech and is useful for many applications. However, it is still challenging to estimate pitch in strong noise. In this paper, we propose a joint training approach to determinate pitch. First, a Bidirectional Long Short-Term Memory Recurrent Neural Networks (BLSTMRNN) is trained to map the noisy to clean speech features. Second, the pitch estimation is also a BLSTM-RNN model. The feature mapping neural network serves as a noise normalization module aiming at explicitly generating the clean features which are easier to estimate pitch by the following neural network. BLSTM-RNN is trained on sequential frame-level features and capable of learning temporal dynamics. We also propose to take into account bottleneck features for pitch estimation. The experimental results show that the proposed method can obtain accurate pitch estimation and they show good generalization ability to new speakers and noisy conditions. The proposed approach also significantly outperforms other state-of-the-art pitch estimation algorithms. Bin Liu 0041, Jianhua Tao 0001, Dawei Zhang 0001, Yibin Zheng |
ICASSP | 1 |
| 2017 | Investigating Efficient Feature Representation Methods and Training Objective for BLSTM-Based Phone Duration Prediction
Yibin Zheng, Jianhua Tao 0001, Zhengqi Wen, Ya Li 0001, Bin Liu 0041 |
INTERSPEECH | 5 |
| 2016 | Extraction of tongue contour in real-time magnetic resonance imaging sequencesabstractReal-time magnetic resonance imaging (rtMRI) is becoming a practical tool in speech production research and language pathology observation. It is still a challenge to extract the tongue contour accurately in rtMRI sequences, since tongue is a soft tissue and often touches other organs such as lips and upper mandible. This paper proposes a novel semiautomatic tongue contour extraction method from rtMRI sequences. The initial boundary image is obtained by combined multi-directional Sobel operators in tongue movement region; then a boundary intensity map is constructed to find the most probable tongue contour points by searching for the optimal boundary route with Viterbi algorithm; finally the tongue contour is obtained using B-Spline approximation. The proposed method could obtain accurate tongue contour from rtMRI sequences, even in the cases that some parts of tongue touch other organs. Experiments demonstrate the robustness of the proposed method. Dawei Zhang 0001, Jianhua Tao 0001, Bin Liu 0041, Danish Bukhari |
ICASSP | 5 |
| 2016 | A Novel Research to Artificial Bandwidth Extension Based on Deep BLSTM Recurrent Neural Networks and Exemplar-Based Sparse Representation
Bin Liu 0041, Jianhua Tao 0001 |
INTERSPEECH | 1 |
| 2015 | Estimate articulatory MRI series from acoustic signal using deep architectureabstractThis paper presents our work on acoustic-to-articulatory inversion mapping, in which, the articulatory data is the MRI series for articulators on mid-sagittal plan. Deep architectures based on restricted Boltzmann machine (RBM) and linear regression are employed to construct the audio-visual mapping. We test two architectures to initialize the neural network: the bottom-up stacked RBM with top regression layer architecture and the one with extra Gaussian-Bernoulli RBM on the top of the former architecture. GMM-based mapping is used as baseline method. The MRI data from USC-TIMIT database is used for the training. The experimental results show that the deep regression network is an effective model to construct the mapping from acoustic speech signal to articulatory MRI series, and also indicate that it is a better strategy to initial the top layer as Gaussian-Bernoulli RBM to compress the MRI data before the liner regression. Hao Li 0078, Jianhua Tao 0001, Bin Liu 0041 |
ICASSP | 4 |
| 2015 | A novel method of artificial bandwidth extension using deep architecture
Bin Liu 0041, Jianhua Tao 0001, Zhengqi Wen, Ya Li 0001, Danish Bukhari |
INTERSPEECH | 1 |
| 2015 | User behavior fusion in dialog management with multi-modal history cues
Jianhua Tao 0001, Linlin Chao, Hao Li 0078, Dawei Zhang 0001, Hao Che, Tingli Gao, Bin Liu 0041 |
Multim. Tools Appl. | 8 |