EDBT 2026 Demo / reviewers in the wild / expert
Mingyue Niu
dblp:231/3883
· DBLP profile ↗
27ranked-venue papers
14as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 8 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 7 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CMTNet: A collaborative mamba-transformer network with spatial-temporal cross-fusion for speech emotion recognition
Shihe Dong, Jiajun Wei, Yibing Zhu, Zhuhong Shao, Mingyue Niu, Xiaohui Tan, Yinan Jiang, Rongyin Qin |
Pattern Recognit. | 7 |
| 2026 | CausalPose: Causal visuo-tactile fusion for robust 6-DoF object pose estimation
Peiliang Wu, Yuanzhi Li, Mingyue Niu, Fengda Zhao, Ziying Song, Yongtao Yang |
Pattern Recognit. | 4 |
| 2026 | Multimodal Local Global Interaction Networks for Automatic Depression Severity EstimationabstractPhysiological studies have shown that differences between depressed and healthy individuals are manifested in the audio and video modalities. Hence, some researchers have combined local and global information from audio or video modality to obtain the unimodal representation. Attention mechanisms or Multi-Layer Perceptrons (MLPs) are then used to complete the fusion of different representations. However, attention mechanisms or MLPs is essentially a linear aggregation manner, and lacks the ability to explore the element-wise interaction between local and global representations within and across modalities, which affects the accuracy of estimating the depression severity. To this end, we propose a Representation Interaction (RI) module, which uses the mutual linear adjustment to achieve element-wise interaction between representations. Thus, the RI module can be seen as an mutual observation of two representations, which helps to achieve complementary advantages and improve the model’s ability to characterize depression cues. Furthermore, since the interaction process generates multiple representations, we propose a Multi-representation Prediction (MP) module. This module implements multi-representation vectorization in a hierarchical manner from summarizing a single representation to aggregating multiple representations, and adopts the attention mechanism to obtain the estimation of an individual depression severity. In this way, we use the RI and MP modules to construct the Multimodal Local Global Interaction (MLGI) network. The experimental performance on AVEC 2013 and AVEC 2014 depression datasets demonstrates the effectiveness of our method. Mingyue Niu, Zhuhong Shao, Yongjun He 0002, Jianhua Tao 0001, Björn W. Schuller |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Facial action units guided graph representation learning for multimodal depression detection
Changzeng Fu, Fengkui Qian, Yikai Su, Kaifeng Su, Siyang Song, Mingyue Niu, Zhigang Liu 0014, Carlos Toshinori Ishi, Hiroshi Ishiguro |
Neurocomputing | 6 |
| 2025 | Examining the Fourier Spectrum of Speech Signal From a Time-Frequency Perspective for Automatic Depression Level PredictionabstractCurrently, many studies use Fourier amplitude spectra of speech signals to predict depression levels. However, those works often treat Fourier amplitude spectra as images or sequences to capture depression cues using convolutional neural networks or multilayer perceptrons. Therefore, they ignore the complex element composition and time-frequency attributes of Fourier spectra, which is not conducive to capturing the differences among individuals with different depression levels. For this reason, we construct a Time-Frequency Self-Embedding (TFSE) module, which not only stores the correlation relationship among real (imaginary) parts of Fourier spectra of different subjects from the time-frequency perspective, but also maintain the physical properties of data through the weight embedding process. Besides, Global Average Pooling (GAP) or linear layers are difficult to balance both temporal and frequency dimensions in the vectorization process. Therefore, we construct a Time-Frequency Tensor Vectorization (TFTV) module, which summarizes each channel along time and frequency dimensions, and then generates the vectorization result by integrating various channels. In this way, we combine TFSE and TFTV modules to form our SpectrumFormer model for predicting depression levels. Evaluation indicators on AVEC 2013 and AVEC 2014 depression databases imply the progressiveness of our model. Mingyue Niu, Jianhua Tao 0001, Yongjun He 0002, Shiqing Zhang, Ming Li 0065 |
IEEE Trans. Affect. Comput. | 1 |
| 2025 | Depression Scale Dictionary Decomposition Framework for Multimodal Automatic Depression Level PredictionabstractCurrently, many researchers aim to achieve automatic depression level prediction via speech and video behavior analysis. However, previous works have struggled to decompose audio and video sequences into the information related to and unrelated to depression scores, hindering the model’s perception of depression cues. Besides, previous works implement multimodal fusion using attention mechanisms or linear layers, but failed to simultaneously consider the Euclidean relationship among tokens and the non-Euclidean relationship among channels, which bring limitations in capturing depression cues. In response to the above issues, we propose a depression scale dictionary decomposition framework, which mainly includes a Bidirectional Dictionary Decomposition (BDD) module and a Bidirectional Multimodal Fusion (BMF) module. The BDD module can use the dictionaries generated based on the depression scale to semantically decompose audio and video sequences into the information related to and unrelated to depression scores along token and channel dimensions for promoting depression cue perception. Moreover, considering the respective characteristics of tokens and channels, the BMF module uses linear layers and graph convolution to achieve cross-modal mixing, which is used to aggregate audio and video sequences for predicting depression levels. The validation on AVEC 2013, AVEC 2014 and DAIC-WOZ datasets demonstrates our method’s superiority. Mingyue Niu, Jibing Gong, Bin Liu 0041, Jianhua Tao 0001, Björn W. Schuller |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | TTFNet: Temporal-Frequency Features Fusion Network for Speech Based Automatic Depression Recognition and AssessmentabstractRelated studies have revealed that the phonological features of depressed patients are different from those of healthy individuals. With the increasing prevalence of depression, an objective and convenient approach for early screening is necessary. To this end, we propose an automatic depression detection method based on hybrid speech features extracted by deep learning, dubbed as TTFNet. Firstly, to effectively excavate the intrinsic relationship among multidimensional dynamic features in the frequency domain, the log-Mel spectrogram of raw speech and its related derivatives are encoded into quaternion representation. Then, the innovatively designed quaternion VisionLSTM is utilized to capture their synergistic effects. Simultaneously, we integrate sLSTM with the pre-trained wav2vec 2.0 model to fully acquire the temporal features. In addition, to further exploit the complementarity between temporal and frequency features, we design an XConformer block for cross-sequence interactions, which ingeniously combines self-attention mechanisms and convolutional modules. Based on this block, the dual-path fusion module closely utilizes the mutual promotion of features from different domains, thereby enhancing generalization capability of the proposed model. Extensive experiments conducted on the AVEC 2013, AVEC 2014, DAIC-WOZ and E-DAIC datasets demonstrate that our method outperforms current state-of-the-art methods in both depression recognition and severity prediction tasks. Xiyuan Chen 0004, Zhuhong Shao, Yinan Jiang, Runsen Chen, Bicao Li, Mingyue Niu, Hongguang Chen, Jiasong Wu |
IEEE J. Biomed. Health Informatics | 7 |
| 2024 | Dense Coordinate Channel Attention Network for Depression Level Estimation from Speech
Ziping Zhao 0001, Shizhao Liu, Mingyue Niu, Haishuai Wang, Björn W. Schuller |
ICPR (13) | 3 |
| 2024 | PointTransform Networks for automatic depression level prediction via facial keypoints
Mingyue Niu, Ming Li 0065, Changzeng Fu |
Knowl. Based Syst. | 1 |
| 2024 | WavDepressionNet: Automatic Depression Level Prediction via Raw Speech SignalsabstractPhysiological reports have confirmed that there are differences in speech signals between depressed and healthy individuals. Therefore, as an application in the field of affective computing, automatic depression level prediction through speech signals has received the attention of researchers, which often estimate the depression severity of individuals by the Fourier or Mel spectrograms of speech signals. However, some studies on speech emotion recognition suggest that directly modeling the raw speech signal is more helpful for extracting emotion-related information. Inspired by this fact, we develop a WavDepressionNet to model raw speech signals for the improvement of prediction accuracy. In our method, a representation block is proposed to find a set of basis vectors to construct the optimal transformation space and generate the transformation result (named Depression Feature Map, DFM) of speech signal for facilitating the perception of depression cues. We further propose an assessment block, which cannot only use the designed spatiotemporal self-calibration mechanism to calibrate the DFM and highlight the useful elements, but also aggregates the calibrated DFM across various temporal ranges with the dilated convolution. Experimental results on the AVEC 2013 and AVEC 2014 depression databases demonstrate the effectiveness of our approach over previous works. Mingyue Niu, Jianhua Tao 0001, Ya Li 0001 |
IEEE Trans. Affect. Comput. | 1 |
| 2024 | MTDAN: A Lightweight Multi-Scale Temporal Difference Attention Networks for Automated Video Depression DetectionabstractDeep learning based video depression analysis has been recently an interesting and challenging topic. Most of existing works focus on learning single-scale facial dynamics of participants for depression detection. Besides, they usually adopt expensive deep learning models with high computational complexity, resulting in difficulty in real-time clinical applications. To address these two issues, this work proposes a lightweight Multi-scale Temporal Difference Attention Networks (MTDAN) integrating the temporal difference and attention mechanism to model both short-term and long-term temporal facial behaviors for automated video depression detection. Initially, two simple yet effective sub-branches, i.e., a Short-term Temporal Difference Attention Network (ST-TDAN), and a Long-term Temporal Difference Attention Network (LT-TDAN), are designed to perform individually short-term and long-term depressive behavior modeling. Then, a simple Interactive Multi-head Attention Fusion (IMHAF) strategy is employed for integrating short-term and long-term spatiotemporal features, followed by a linear fully-collected layer for depression score prediction. Experiments on two public AVEC2013 and AVEC2014 datasets show that our proposed method not only achieves highly competitive performance to state-of-the-art methods, but also has much smaller computational complexity than them on video depression detection tasks. Shiqing Zhang, Xingnan Zhang, Xiaoming Zhao 0002, Jiangxiong Fang, Mingyue Niu, Ziping Zhao 0001, Jun Yu 0002, Qi Tian 0001 |
IEEE Trans. Affect. Comput. | 5 |
| 2024 | DepressionMLP: A Multi-Layer Perceptron Architecture for Automatic Depression Level Prediction via Facial Keypoints and Action UnitsabstractPhysiological studies have confirmed that there are differences in facial activities between depressed and healthy individuals. Therefore, while protecting the privacy of subjects, substantial efforts are made to predict the depression severity of individuals by analyzing Facial Keypoints Representation Sequences (FKRS) and Action Units Representation Sequences (AURS). However, those works has struggled to examine the spatial distribution and temporal changes of Facial Keypoints (FKs) and Action Units (AUs) simultaneously, which is limited in extracting the facial dynamics characterizing depressive cues. Besides, those works don’t realize the complementarity of effective information extracted from FKRS and AURS, which reduces the prediction accuracy. To this end, we intend to use the recently proposed Multi-Layer Perceptrons with gating (gMLP) architecture to process FKRS and AURS for predicting depression levels. However, the channel projection in the gMLP disrupts the spatial distribution of FKs and AUs, leading to input and output sequences not having the same spatiotemporal attributes. This discrepancy hinders the additivity of residual connections in a physical sense. Therefore, we construct a novel MLP architecture named DepressionMLP. In this model, we propose the Dual Gating (DG) and Mutual Guidance (MG) modules. The DG module embeds cross-location and cross-frame gating results into the input sequence to maintain the physical properties of data to make up for the shortcomings of gMLP. The MG module takes the global information of FKRS (AURS) as a guidance mask to filter the AURS (FKRS) to achieve the interaction between FKRS and AURS. Experimental results on several benchmark datasets show the effectiveness of our method. Mingyue Niu, Ya Li 0001, Jianhua Tao 0001, Xiuzhuang Zhou, Björn W. Schuller |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Multimodal Spatiotemporal Representation for Automatic Depression Level DetectionabstractPhysiological studies have shown that there are some differences in speech and facial activities between depressive and healthy individuals. Based on this fact, we propose a novel spatio-temporal attention (STA) network and a multimodal attention feature fusion (MAFF) strategy to obtain the multimodal representation of depression cues for predicting the individual depression level. Specifically, we first divide the speech amplitude spectrum/video into fixed-length segments and input these segments into the STA network, which not only integrates the spatial and temporal information through attention mechanism, but also emphasizes the audio/video frames related to depression detection. The audio/video segment-level feature is obtained from the output of the last full connection layer of the STA network. Second, this article employs the eigen evolution pooling method to summarize the changes of each dimension of the audio/video segment-level features to aggregate them into the audio/video level feature. Third, the multimodal representation with modal complementary information is generated using the MAFF and inputs into the support vector regression predictor for estimating depression severity. Experimental results on the AVEC2013 and AVEC2014 depression databases illustrate the effectiveness of our method. Mingyue Niu, Jianhua Tao 0001, Bin Liu 0041, Jian Huang 0014, Zheng Lian 0004 |
IEEE Trans. Affect. Comput. | 1 |
| 2023 | Dual Attention and Element Recalibration Networks for Automatic Depression Level PredictionabstractPhysiological studies have identified that facial dynamics can be considered as biomarkers to analyze depression severity. This paper accordingly develops a Dual Attention and Element Recalibration (DAER) network to extract facial changes to predict the depression level. In this model, we propose two blocks: a Dual Attention (DA) block and Element Recalibration (ER) block. The DA block uses the self-attention to investigate the dynamic changes in the representation sequence of a facial video segment. It further examines the influence of feature components of the representation sequence on depression level prediction through bilinear-attention. Moreover, to improve the representation ability of network, the ER block is used to obtain the global information to recalibrate each element of the tensor. Adopting this approach, for the depression level prediction task, we first divide the long-term video into fixed-length segments and use the trained ResNet50 to encode each frame to generate the representation sequences of video segments. Second, the representation sequences are input into DAER network to obtain the depression level scores. Finally, the average of these scores yields the prediction result corresponding to the long-term video. Experiments on publicly available AVEC 2013 and AVEC 2014 depression databases illustrate the effectiveness of our method. Mingyue Niu, Ziping Zhao 0001, Jianhua Tao 0001, Ya Li 0001, Björn W. Schuller |
IEEE Trans. Affect. Comput. | 1 |
| 2022 | Csenet: Complex Squeeze-and-Excitation Network for Speech Depression Level PredictionabstractAutomatic speech depression level prediction (SDLP) is a very challenging problem in affective computing. There are many studies that have acquired quite good performances for SDLP. However, most of the input speech features of these studies are based on the amplitude spectrogram, which loses the phase spectrogram information. Therefore, these speech features may lose some important information related to depression. In order to make full use of speech information, this paper proposes a complex squeeze-and-excitation network (CSENet) for SDLP. The complex spectrogram is used as the input speech feature, which contains both amplitude and phase spectrogram. In addition, to acquire a discriminative feature, the squeeze-and-excitation residual network is employed to extract deep speech feature. Finally, the attentive temporal pooling is utilized to dynamically select more important information according to the attention mechanisms. Experimental results on the AVEC 2013 and AVEC 2014 datasets prove the effectiveness of our proposed method. As for the mean absolute error (MAE) evaluation metric on AVEC 2013, our proposed method acquires state-of-the-art performance. Cunhang Fan, Zhao Lv, Shengbing Pei, Mingyue Niu |
ICASSP | 4 |
| 2022 | Automatic Depression Level Assessment from Speech By Long-Term Global Information EmbeddingabstractDepression is a serious mood disorder which brings negative effects on people's social activities. Therefore, growing attention has been paid to automatic depression assessment, especially from speech. However, most of the previous work uses hand-crafted features or deep neural network-based feature extractors to obtain deep features and then feed them into a classifier or a regression, which ignores the temporal relation of these features. To address this issue, this paper proposes a global information embedding (GIE) to make use of the long-term global information of depression and re-weight the LSTM output sequence. The short-term features are then pooled into long-term features by LASSO optimization to further improve the accuracy of depression recognition. Experiments on AVEC 2013 and AVEC 2014 verified the proposed method, and the RMSEs are 9.63 and 9.40, respectively. Ya Li 0001, Mingyue Niu, Ziping Zhao 0001, Jianhua Tao 0001 |
ICASSP | 2 |
| 2022 | Automatic Respiratory Sound Classification Via Multi-Branch Temporal Convolutional NetworkabstractAutomated classification of respiratory sounds has become an active research area in recent years. While recent studies have utilised deep learning methods to aid with respiratory sound classification, the performance is heavily influenced by the datasets available for respiratory sound classification tasks, which tend to be smaller and imbalanced. In this paper, we propose to explore the effectiveness of a multi-branch Temporal Convolutional Network (TCN) architecture integrated with Squeeze-and-Excitation Network (SEnet), a system denoted herein as MBTCNSE, for respiratory sound classification. To the best of the authors’ knowledge, this is the first time that such a hybrid architecture has been employed for respiratory sounds classification. Experiments based on the ICBHI challenge respiratory sound dataset demonstrate the effectiveness of our method. Ziping Zhao 0001, Mingyue Niu, Haishuai Wang, Zixing Zhang 0001, Ya Li 0001 |
ICASSP | 3 |
| 2022 | Depressioner: Facial dynamic representation for automatic depression level prediction
Mingyue Niu, Ya Li 0001, Bin Liu 0041 |
Expert Syst. Appl. | 1 |
| 2022 | Selective Element and Two Orders Vectorization Networks for Automatic Depression Severity Diagnosis via Facial ChangesabstractPhysiological studies have shown that healthy and depressed individuals present different facial changes. Thus, many researchers have attempted to use Convolutional Neural Networks (CNNs) to extract high-level facial dynamic representations for predicting depression severity. However, the max-pooling (or average-pooling) layers in the CNN lead to the loss of subtle depression cues. Without pooling layers, the CNN cannot extract multi-scale information and has difficulties for tensor vectorization. To this end, we propose a Selective Element and Two Orders Vectorization (SE-TOV) network. For the SE-TOV network, an SE block is constructed to adaptively select the effective elements from the tensors obtained by receptive fields of different sizes. Moreover, we propose a TOV block for vectorizing a high-dimensional tensor. On the one hand, TOV block inputs a tensor into the Global Average Pooling layer to obtain the first-order vectorization result. On the other hand, it takes principal components of the correlation matrix of channels in a tensor as the second-order vectorization result. Experimental results on AVEC 2013 (RMSE$=7.42$, MAE$=6.09$) and AVEC 2014 (RMSE$=7.39$, MAE$=5.87$) depression databases illustrate the superiority of our approach over previous works. Mingyue Niu, Ziping Zhao 0001, Jianhua Tao 0001, Ya Li 0001, Björn W. Schuller |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Multi-Scale and Multi-Region Facial Discriminative Representation for Automatic Depression Level PredictionabstractPhysiological studies have shown that differences in facial activities between depressed patients and normal individuals are manifested in different local facial regions and the durations of these activities are not the same. But most previous works extract features from the entire facial region at a fixed time scale to predict the individual depression level. Thus, they are inadequate in capturing dynamic facial changes. For these reasons, we propose a multi-scale and multi-region fa-cial dynamic representation method to improve the prediction performance. In particular, we firstly use multiple time scales to divide the original long-term video into segments containing different facial regions. Secondly, the segment-level feature is extracted by 3D convolution neural network to characterize the facial activities with different durations in different facial regions. Thirdly, this paper adopts eigen evolution pooling and gradient boosting decision tree to aggregate these segment-level features and select discriminative elements to generate the video-level feature. Finally, the depression level is predicted using support vector regression. Experiments are conducted on AVEC2013 and AVEC2014. The results demonstrate that our method achieves better performance than the previous works. Mingyue Niu, Jianhua Tao 0001, Bin Liu 0041 |
ICASSP | 1 |
| 2021 | TDCA-Net: Time-Domain Channel Attention Network for Depression Detection
Cong Cai, Mingyue Niu, Bin Liu 0041, Jianhua Tao 0001, Xuefei Liu |
Interspeech | 2 |
| 2021 | A time-frequency channel attention and vectorization network for automatic depression level prediction
Mingyue Niu, Bin Liu 0041, Jianhua Tao 0001 |
Neurocomputing | 1 |
| 2020 | ParamE: Regarding Neural Network Parameters as Relation Embeddings for Knowledge Graph CompletionabstractWe study the task of learning entity and relation embeddings in knowledge graphs for predicting missing links. Previous translational models on link prediction make use of translational properties but lack enough expressiveness, while the convolution neural network based model (ConvE) takes advantage of the great nonlinearity fitting ability of neural networks but overlooks translational properties. In this paper, we propose a new knowledge graph embedding model called ParamE which can utilize the two advantages together. In ParamE, head entity embeddings, relation embeddings and tail entity embeddings are regarded as the input, parameters and output of a neural network respectively. Since parameters in networks are effective in converting input to output, taking neural network parameters as relation embeddings makes ParamE much more expressive and translational. In addition, the entity and relation embeddings in ParamE are from feature space and parameter space respectively, which is in line with the essence that entities and relations are supposed to be mapped into two different spaces. We evaluate the performances of ParamE on standard FB15k-237 and WN18RR datasets, and experiments show ParamE can significantly outperform existing state-of-the-art models, such as ConvE, SACN, RotatE and D4-STE/Gumbel. Feihu Che, Dawei Zhang 0001, Jianhua Tao 0001, Mingyue Niu, Bocheng Zhao |
AAAI | 4 |
| 2020 | Multimodal Transformer Fusion for Continuous Emotion RecognitionabstractMultimodal fusion increases the performance of emotion recognition because of the complementarity of different modalities. Compared with decision level and feature level fusion, model level fusion makes better use of the advantages of deep neural networks. In this work, we utilize the Transformer model to fuse audio-visual modalities on the model level. Specifically, the multi-head attention produces multimodal emotional intermediate representations from common semantic feature space after encoding audio and visual modalities. Meanwhile, it also can learn long-term temporal dependencies with self-attention mechanism effectively. The experiments, on the AVEC 2017 database, shows the superiority of model level fusion than other fusion strategies. Moreover, we combine the Transformer model and LSTM to further improve the performance, which achieves better results than other methods. Jian Huang 0014, Jianhua Tao 0001, Bin Liu 0041, Zheng Lian 0004, Mingyue Niu |
ICASSP | 5 |
| 2019 | Efficient Modeling of Long Temporal Contexts for Continuous Emotion RecognitionabstractContinuous emotion recognition is a challenging task due to its difficulty in modeling long-term contexts dependencies. Prior researches have exploited emotional temporal contexts from two perspectives, which are based on feature representations and emotional models. In this paper, we explore the model based approaches for continuous emotion recognition. Specifically, three temporal models including LSTM, TDNN and multi-head attention models are utilized to learn long-term contexts dependencies based on short-term feature representations. The temporal information learned by the temporal models allows the network to more easily exploit the slow changing dynamics between emotional states. Our experimental results demonstrate that the temporal models can model emotional long-term dynamic information effectively. Multi-head attention model achieves best performance among three models and multi-model combination models further improve the performance of continuous emotion recognition significantly. Jian Huang 0014, Jianhua Tao 0001, Bin Liu 0041, Zhen Lian, Mingyue Niu |
ACII | 5 |
| 2019 | Discriminative Video Representation with Temporal Order for Micro-expression RecognitionabstractMicro-expression recognition is a challenging task due to its low intensity and short duration and how to extract the subtle facial changes is a key issue in this field. Although there are many methods attempt to cope with this problem, they are difficult to encode the temporal order of all frames in the video clips. For these reasons, this paper employs rank pooling and ℓ2,1-norm to obtain the discriminative video representation with temporal order. In particular, we extract Local Two-Order Gradient Pattern (LTOGP) feature of each frame to describe the subtle information. Then, the video representation is generated by using rank pooling, which captures the temporal order among all frames. Furthermore, considering the sparsity of ℓ2,1-norm, we can select those discriminant features. Finally, micro-expression classification is accomplished using SVM. Experiments are conducted on two publicly available micro-expression databases i.e. CASME and CASME2. The results demonstrate that our method achieves better performance than the state-of-the-art algorithms. Mingyue Niu, Jianhua Tao 0001, Ya Li 0001, Jian Huang 0014, Zheng Lian 0004 |
ICASSP | 1 |
| 2019 | Automatic Depression Level Detection via ℓp-Norm Pooling
Mingyue Niu, Jianhua Tao 0001, Bin Liu 0041, Cunhang Fan |
INTERSPEECH | 1 |