EDBT 2026 Demo / reviewers in the wild / expert
Haifeng Li 0001
dblp:17/246-1
· DBLP profile ↗
32ranked-venue papers
3as first author
19since 2021 · last 2026
0000-0002-2534-2299ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 7 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Leveraging Visual Blur Perception Characteristics for EEG DecodingabstractIn recent years, electroencephalography (EEG)-based visual decoding research has become a key direction for revealing brain processing mechanisms and realizing brain-computer interfaces. This emerging field has attracted extensive attention in the fields of brain science, cognitive neuroscience, and artificial intelligence. Among various approaches, contrastive learning has demonstrated strong performance in aligning multi-modal data, effectively enabling unified representations across modalities. However, during human visual perception, images are often subject to varying degrees of blurring due to the uneven distribution of retinal photoreceptor cells and the limited speed of lens accommodation. To address the mismatch between EEG and visual representations, we propose a novel visual decoding framework inspired by human perceptual blurring. Specifically, multi-level Gaussian blurring is applied to the visual stimuli to simulate human visual characteristics, followed by a feature selection module to construct robust visual representations. For EEG decoding, we design a lightweight and efficient network employing positively constrained spatial convolutions to identify channels associated with visual processing. The EEG and visual features are then aligned using contrastive learning. We evaluate the proposed framework on the Things-EEG dataset. Experimental results show significant improvements in the zero-shot brain-to-image retrieval task, achieving a top-1 accuracy of 80% and a top-5 accuracy of 96.9%, surpassing previous state-of-the-art methods by margins of 29.1% and 17.2%, respectively. These findings highlight the potential of incorporating perceptual properties into EEG-based visual decoding. Wenchao Liu 0004, Hongwei Li 0024, Zhouyang Xu, Lin Ma 0003, Haifeng Li 0001 |
AAAI | 5 |
| 2026 | Modeling the evolution dynamics to enhance micro-expression recognition
Wenchao Liu 0004, Lin Ma 0003, Haifeng Li 0001 |
Frontiers Comput. Sci. | 5 |
| 2025 | Research of Decoding Cognitive Processes Based on Multimodal Brain Signal FusionabstractWith the advancement of neuroscience and artificial intelligence, there is a growing demand for decoding the brain's complex cognitive activities. However, the analysis of any single signal modality is inherently limited by the complexity of brain function. This study aims to decode memory-related cognitive processes by integrating two complementary brain signal modalities: intracranial electroencephalography (iEEG) and functional magnetic resonance imaging (fMRI). We present a complete research framework that includes data preprocessing, cross-modal feature mapping, functional connectivity analysis, and a neural network model for cognitive state classification. Specifically, we propose a multimodal fusion model for cognitive state analysis that utilizes an attention mechanism to fuse iEEG spectral features with fMRI-BOLD representations. This model is used to classify whether a subject is actively engaged in a memory recognition task. Experiments conducted on a multimodal signal dataset confirm that our proposed model achieves a significant improvement in classification accuracy compared to other models. This research contributes a new analytical method for multimodal brain signal studies and provides new insights into the neural mechanisms of memory and the field of braincomputer interfaces. Shihang Ding, Kaichen Lan, Lin Ma 0003, Haifeng Li 0001 |
BIBM | 4 |
| 2025 | Real-Time Feedback-Based Motor Imagery EEG Acquisition Method Using Brain Foundation ModelabstractMotor imagery (MI) is widely used in brain-computer interfaces (BCIs), but conventional EEG acquisition often produces low-quality signals. Feedback methods improve EEG quality, but typically need initial session data to train classifiers. This increases complexity and delays feedback. We solve this using a brain foundation model (EEGPT), which leverages self-supervised pre-training on large-scale MI data to enable accurate classification with minimal new samples. Our method combines EEGPT's encoder (frozen) for feature extraction with an online updated Support Vector Machine (SVM) classifier. This provides real-time feedback during EEG acquisition. Our approach shows superior efficiency and accuracy under limited data versus models trained from scratch on BNCI2014 and ZHOU2016 datasets. In a 5 -class MI experiment with 7 participants, data acquired with our feedback method showed 4.1% higher average accuracy across multiple models. Our method processes EEG quickly (inference:$9.92 \pm 0.51 ~\text{ms}$per sample; SVM training:$0.276 \pm 0.008 ~\mathrm{s})$, enabling early high-quality feedback with minimal calibration. This significantly improves BCI efficiency and practical deployment. Lin Ma 0003, Haifeng Li 0001 |
BIBM | 3 |
| 2025 | MIBM: Interpreting EEGNet via MODET-Based Interaction Backtracking MethodabstractWith the rapid development of deep neural networks, significant progress has been made in the analysis of physiological signals such as EEG, greatly improving classification performance. However, the inherent “black-box” nature of these models limits our ability to understand how different brain regions coordinate to drive decisions-especially given that modern EEG architectures (e.g., EEGNet) use gating mechanisms that activate distinct computational pathways for different subjects. To address this, we introduce the MODET-based Interaction Backtracking Method (MIBM), which converts a standard EEGNet into a shallow, wide, and intrinsically interpretable Multi-Order Descartes Expansion Neural Network (MODENN). This “deep-to-broad” transformation makes high-order feature interactions explicit and traceable. Experimentally, the MIBMtransformed EEGNet preserves-or even surpasses-the original model's accuracy across subjects. By exposing each subject's gated pathways as a clear computation graph, MIBM substantially enhances transparency and trustworthiness, marking a paradigm shift from traditional single-feature attribution toward a holistic, network-level interpretability aligned with systems neuroscience. Hongjia Zhu, Zexi Liu, Lin Ma 0003, Haifeng Li 0001 |
BIBM | 5 |
| 2025 | TimeStacker: A Novel Framework with Multilevel Observation for Capturing Nonstationary Patterns in Time Series ForecastingabstractReal-world time series inherently exhibit significant non-stationarity, posing substantial challenges for forecasting. To address this issue, this paper proposes a novel prediction framework, TimeStacker, designed to overcome the limitations of existing models in capturing the characteristics of non-stationary signals. By employing a unique stacking mechanism, TimeStacker effectively captures global signal features while thoroughly exploring local details. Furthermore, the framework integrates a frequency-based self-attention module, significantly enhancing its feature modeling capabilities. Experimental results demonstrate that TimeStacker achieves outstanding performance across multiple real-world datasets, including those from the energy, finance, and weather domains. It not only delivers superior predictive accuracy but also exhibits remarkable advantages with fewer parameters and higher computational efficiency. Qinglong Liu, Cong Xu 0004, Lin Ma 0003, Haifeng Li 0001 |
ICML | 6 |
| 2025 | Deep-to-broad technique to accelerate deep neural networks
Hongjia Zhu, Cong Xu 0004, Lin Ma 0003, Haifeng Li 0001 |
Expert Syst. Appl. | 4 |
| 2024 | fNIRS emotion recognition method based on multi-scale fusion featuresabstractFunctional near-infrared imaging is a brain imaging technology that measures changes in cerebral blood oxygenation. Due to its high portability and spatial resolution, fNIRS is widely employed in emotion cognition. However, current feature extraction methods for fNIRS predominantly rely on statistical feature engineering, often lacking deeper exploration and analysis of both long-term and short-term characteristics of fNIRS signals. In response to this limitation, this paper presents research conducted on the fNIRS emotion dataset NEMO and introduces an emotion recognition method based on multi-scale fusion features. By integrating vector features and time-delay embedding features, the proposed method facilitates the identification of dimensional emotions. The model demonstrates its efficacy by effectively distinguishing emotions in binary classifications of valence and arousal, as well as in the four-class classification of dimensional emotions. These results confirm the utility of fNIRS in emotion decoding and emotional state assessment, offering valuable blood oxygen signal features for emotion recognition. Furthermore, this work provides a viable approach for decoding the neural activities associated with dimensional emotions, and introduces a new perspective for advancing our understanding of brain processes related to emotion. Shihang Ding, Cong Xu 0004, Hongjian Bo, Lin Ma 0003, Haifeng Li 0001 |
BIBM | 7 |
| 2024 | Multimodal Brain Signal Analysis with State-Space Modeling: A Study on Working MemoryabstractWorking memory (WM) is a fundamental cognitive function intricately linked to various neurological and psychiatric conditions. In this study, we introduce Neural-SSM, a novel computational model based on the State Space Model (SSM), designed to effectively capture and characterize neural activity patterns. The model was trained on multimodal brain signals, specifically electroencephalogram (EEG) and functional near-infrared spectroscopy (fNIRS) data collected during the N-back tasks. Experimental results indicate that the Embedding Unit achieved an accuracy of 81.41% in classifying N-back task difficulty, while the Attention GRU component demonstrated precise regression of hemodynamic responses with a mean squared error (MSE) of 0.1289. Analysis of the attention matrix revealed enhanced directional connectivity from EEG to fNIRS channels and identified long-range pathways from the occipital to frontal cortex, which intensified with increasing task difficulty. These findings highlight the model’s potential in advancing the understanding of brain activities in cognitive processes, with significant implications for biomedical engineering and neuroimaging research. Qinglong Liu, Shihang Ding, Hongjian Bo, Cong Xu 0004, Lin Ma 0003, Haifeng Li 0001 |
BIBM | 8 |
| 2024 | Exploring the Neural Dynamics in Temporal Lobe Epilepsy: A Study using Transformer and Hidden Markov ModelsabstractAdvancing the understanding of temporal lobe epilepsy (TLE) requires sophisticated analytical tools. In this study, we introduce a hybrid model, namely the HMM-Wavformer, aiming at identifying the phasic brain activity patterns during seizures on Stereo-electroencephalography (SEEG) records. The model is composed of a wavelet packet decomposition (WPD) based signal processing module, an embedding module for spatial feature extraction, and a Transformer module to weigh the time-series frequency importance. The model is trained with a downstream seizure detection task on the HUP-iEEG dataset, demonstrating an accuracy of 92.75%. Frequency analysis identifies the most sensitive bands in TLE seizure detection. The Hidden Markov Model (HMM) is applied for the time-series analysis, categorizing the seizures into three ictal phases. Complementary analyses using power spectra and brain networks pinpoint biomarkers for each phase. The analysis results indicate that, the HMM-Wavformer model is able to effectively depict the neural dynamics of TLE seizures, aligning with prior medical studies, and provide a more detailed description of the staged characteristics of these seizures. Zhiguo Lin, Shihang Ding, Chunying Fang, Hongjian Bo, Cong Xu 0004, Shengkun Yu, Yifei Gu, Tiejun Zhao, Haifeng Li 0001 |
BIBM | 12 |
| 2024 | MCAHNN: Multi-Channel EEG Emotion Recognition Using Attention Mechanism Based on Householder ReflectionabstractEmotions are integral to human cognition, exerting a profound influence on physiological responses, cognitive processes, and decision-making capabilities. Electroencephalography (EEG)-based emotion classification provides a significant methodological approach for the exploration of emotional states. Despite its potential, most current methodologies face challenges in delineating the representational patterns across different brain regions and in effectively classifying emotions from EEG signals. In response, a novel model for emotion recognition is proposed in this paper, which utilizes a multi-channel attention mechanism, designated as MCAHNN. This model incorporates Householder Reflection to enhance the attention mechanism, facilitating the extraction of inter-channel EEG features and simulating inter-regional brain dynamics. Furthermore, 1D convolution is employed to analyze intra-channel relationships. The proposed model has been evaluated on the publicly available DEAP dataset and further tested on the SEED dataset. Experimental results confirm that the MCAHNN model achieves state-of-the-art performance, demonstrating its effectiveness in classifying emotions within multi-center datasets.Code is publicly available at https://github.com/Oreoreoreor/MCAHNN. Qinglong Liu, Shihang Ding, Hongjian Bo, Cong Xu 0004, Lin Ma 0003, Haifeng Li 0001 |
ECAI | 8 |
| 2024 | A Study of Prototypical Network Techniques for Cross-Subject EEG AnalysisabstractElectroencephalography (EEG), a technique that uses electrodes on the scalp to record changes in cerebral electrical potentials, plays a key role in various brain-computer interface applications. However, due to the large differences in EEG distribution across subjects, how to accurately classify cross-subject EEG is still a challenge. In practice, it is unavoidable to collect data for each subject to perform calibration, which is cumbersome and expensive. In this paper, we propose a Cross Subject EEG Prototypical Network (CS-EEGPNet) for accurate few-shot EEG classification. To the best of our knowledge, we are the first to eliminate inter-subject distribution shift in few-shot EEG classification through multiple feature layers computation. Specifically, to avoid the effect of individual differences, we propose a few-shot feature normalization (FSFN) method, which uses the mean and variance of the same subject’s support set features to normalize the query set features. Meanwhile, we design a prototypical module based on auxiliary factors. This module utilizes learnable vectors to store category-related information for assisting the construction of prototypes, thus mitigating the effects of EEG’s randomness and non-stationary characteristics on cross-subject few shot classification. To validate the effectiveness of the method, we selecte the motor imagery paradigm, which is relatively complex in brain-computer interfaces, and test our method on the BCIC IV 2a and HGD datasets. The results demonstrate that our method significantly enhances EEG classification performance, achieving the state-of-the-art accuracies. Wenchao Liu 0004, Guagnyu Wang, Hongjian Bo, Lin Ma 0003, Haifeng Li 0001 |
ECAI | 6 |
| 2024 | Enhancing Micro-Expression Analysis Performance by Effectively Addressing Data ImbalanceabstractMicro-expressions (MEs) are involuntary and quickly displayed facial expressions that reveal subtle psychological activities. Most previous research typically focused on two separate tasks: micro-expression spotting and recognition. We aim to propose a high-precision "spotting+recognition" method that can spot ME intervals from long videos and recognize their emotional categories. Due to the occurrence sparsity of MEs, there is a significant imbalance between the number of micro-expression intervals and non-micro-expression intervals in long videos. This imbalance makes it challenging for models trained using conventional strategies to distinguish true MEs from noise samples caused by head movements, blinking, and macro-expressions, resulting in a high false-positive-rate and reducing the overall performance. We reduce the number of smooth segments to alter the data distribution within the non-micro-expression (non-ME) category. This adjustment enables the model to focus more on the subtle differences between noise samples and ME samples. To achieve this, we design an ingenious training data preparation strategy: using false positive samples from the initial spotting results as non-ME category samples, and using true positive and false negative samples from the initial spotting as emotion category samples. These are combined as the training data, creating a recognition model capable of both emotion classification and non-ME category determination. Additionally, we propose a three-stage micro-expression analysis method, including ME spotting, ME recognition and non-ME intervals removal module. Our method is validated through five-fold cross-validation experiments on the CAS(ME)² and SAMM Long Video datasets, achieving a overall STRS metric of 0.16, which significantly outperformed baseline methods and demonstrated the effectiveness of our approach. Wenchao Liu 0004, Lin Ma 0003, Haifeng Li 0001 |
ACM Multimedia | 5 |
| 2024 | EEGPT: Pretrained Transformer for Universal and Reliable Representation of EEG SignalsabstractElectroencephalography (EEG) is crucial for recording brain activity,
with applications in medicine, neuroscience, and brain-computer interfaces (BCI).
However, challenges such as low signal-to-noise ratio (SNR), high inter-subject variability, and channel mismatch complicate the extraction of robust,
universal EEG representations.
We propose EEGPT, a novel 10-million-parameter pretrained transformer model designed for universal EEG feature extraction.
In EEGPT, a mask-based dual self-supervised learning method for efficient feature extraction is designed.
Compared to other mask-based self-supervised learning methods,
EEGPT introduces spatio-temporal representation alignment.
This involves constructing a self-supervised task based on
EEG representations that possess high SNR and rich semantic information,
rather than on raw signals.
Consequently, this approach mitigates the issue of poor feature quality typically
extracted from low SNR signals.
Additionally, EEGPT's hierarchical structure processes spatial and temporal information separately,
reducing computational complexity while increasing flexibility and adaptability for BCI applications.
By training on a large mixed multi-task EEG dataset, we fully exploit EEGPT's capabilities.
The experiment validates the efficacy and scalability of EEGPT,
achieving state-of-the-art performance on a range of downstream tasks with linear-probing.
Our research advances EEG representation learning, offering innovative solutions for bio-signal processing and AI applications.
The code for this paper is available at: https://github.com/BINE022/EEGPT Wenchao Liu 0004, Cong Xu 0004, Lin Ma 0003, Haifeng Li 0001 |
NeurIPS | 6 |
| 2024 | A novel capsule network based on Multi-Order Descartes Extension Transformation
Hongjia Zhu, Cong Xu 0004, Lin Ma 0003, Haifeng Li 0001 |
Neurocomputing | 4 |
| 2024 | Adaptively Optimized Masking EMD for Separating Intrinsic Oscillatory Modes of Nonstationary SignalsabstractEmpirical mode decomposition (EMD) is a well-established technique to decompose nonstationary signals into intrinsic oscillatory modes, which is essential to catch crucial information hidden in nonstationary signals. However, when decomposing signals with intermittency or closely spaced frequency modes, EMD suffers from severe mode mixing problems (MMP). Therefore, we develop a novel adaptively optimized masking EMD (AOMEMD), a generalized, automated method to solve MMP. The work focuses on two aspects: i) Masking signals with different scales are adaptively iteratively generated based on Hilbert analytic solution and applied to raw data to assist EMD in solving MMP. ii) A universal mode mixing measure index (UMI) is proposed to quantitatively evaluate the extent of mode mixing. An optimization routine is then developed to optimize the amplitude and frequency of masking signals with the objective of minimizing the UMI to further improve the frequency separation performance of AOMEMD. Experiments on synthetic and real-world signals demonstrate that AOMEMD outperforms EMD and its variants in effectively addressing MMP, yielding more reasonable and reliable results. AOMEMD can separate modes with frequency ratios up to 0.9. Congshan Sun, Hongwei Li 0024, Cong Xu 0004, Lin Ma 0003, Haifeng Li 0001 |
IEEE Signal Process. Lett. | 5 |
| 2023 | Micro-Expression Spotting Method Based on AU PrototypeabstractMicro-expressions (MEs) are brief, involuntary facial expressions that reveal genuine emotions, making their accurate detection crucial in various applications, such as security, psychology, and human-computer interaction. Due to its small intensity and short duration, how accurately capturing the subtle movements of micro-expression is a challenging problem. This paper presents a novel AU prototype-based method for micro-expression spotting, which offers high accuracy and robustness. Action Units (AUs) are basic facial actions, such as brow lower and lip corner puller, that are widely used for micro-expression analysis, and an expression can be encoded as a sequence of AUs. Our approach involves designing AU prototypes that record representative dynamic information of AUs. We then calculate the prototype matching index between AU prototypes and the image sequence to construct time-domain prototype matching curves for ME spotting. In the experimental section, AU prototypes derived from CASMEII dataset enable a more intuitive analysis of AU within micro-expressions. Results on the CAS(ME)2 dataset demonstrate that our ME spotting method significantly outperforms existing approaches. This makes our method highly valuable for various application scenarios, potentially enhancing emotion recognition and analysis in real-world settings. Lin Ma 0003, Haifeng Li 0001 |
ECAI | 4 |
| 2022 | MODENN: A Shallow Broad Neural Network Model Based on Multi-Order Descartes ExpansionabstractDeep neural networks have achieved great success in almost every field of artificial intelligence. However, several weaknesses keep bothering researchers due to its hierarchical structure, particularly when large-scale parallelism, faster learning, better performance, and high reliability are required. Inspired by the parallel and large-scale information processing structures in the human brain, a shallow broad neural network model is proposed on a specially designed multi-order Descartes expansion operation. Such Descartes expansion acts as an efficient feature extraction method for the network, improve the separability of the original pattern by transforming the raw data pattern into a high-dimensional feature space, the multi-order Descartes expansion space. As a result, a single-layer perceptron network will be able to accomplish the classification task. The multi-order Descartes expansion neural network (MODENN) is thus created by combining the multi-order Descartes expansion operation and the single-layer perceptron together, and its capacity is proved equivalent to the traditional multi-layer perceptron and the deep neural networks. Three kinds of experiments were implemented, the results showed that the proposed MODENN model retains great potentiality in many aspects, including implementability, parallelizability, performance, robustness, and interpretability, indicating MODENN would be an excellent alternative to mainstream neural networks. Haifeng Li 0001, Cong Xu 0004, Lin Ma 0003, Hongjian Bo, David Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Micro-expression spotting based on optical flow features
Zhongliang Xu, Lin Ma 0003, Haifeng Li 0001 |
Pattern Recognit. Lett. | 4 |
| 2019 | Voice conversion with SI-DNN and KL divergence based mapping without parallel training data
Fenglong Xie, Frank K. Soong, Haifeng Li 0001 |
Speech Commun. | 3 |
| 2018 | Towards Temporal Modelling of Categorical Speech Emotion RecognitionabstractTo model the categorical speech emotion recognition task in a temporal manner, the first challenge arising is how to transfer the categorical label for each utterance into a label sequence.To settle this, we make a hypothesis that an utterance is consisting of emotional and non-emotional segments, and these non-emotional segments correspond to silent regions, short pauses, transitions between phonemes, unvoiced phonemes, etc.With this hypothesis, we propose to treat an utterance's label sequence as a chain of two states: the emotional state denoting the emotional frame and Null denoting the non-emotional frame.Then, we exploit a recurrent neural network based connectionist temporal classification model to automatically label and align an utterance's emotional segments with emotional labels, while non-emotional segments with Nulls.Experimental results on the IEMOCAP corpus validate our hypothesis and also demonstrate the effectiveness of our proposed method compared to the state-of-the-art algorithms. Wenjing Han, Huabin Ruan, Haifeng Li 0001, Björn W. Schuller |
INTERSPEECH | 5 |
| 2016 | A KL divergence and DNN approach to cross-lingual TTSabstractWe propose a Kullback-Leibler divergence (KLD) and deep neural net (DNN) based approach to cross-lingual TTS (CL-TTS) training. A speaker independent DNN (SI-DNN) ASR is used to equalize the speaker difference between a source speaker in L1 and a reference speaker in L2. Two speaker dependent GMM-HMM parametric TTS systems are first trained in the respective languages. The senones sets of the two TTS are matched in the SI-DNN ASR in terms of their output posteriors distributions in KLD. The minimum KLD criterion is used to transform the senones in the source speaker's TTS (L1) to the corresponding "closest" senones in the target language (L2). The new CL-TTS thus trained has been shown to achieve high speaker similarity to the source speaker in L1 while high intelligibility and naturalness are preserved. For untranscribed source speaker's recordings, say, conversational speech, a frame mapping, instead of "senone mapping" is also proposed to achieve a high but slightly inferior CL-TTS. Fenglong Xie, Frank K. Soong, Haifeng Li 0001 |
ICASSP | 3 |
| 2016 | A KL Divergence and DNN-Based Approach to Voice Conversion without Parallel Training Sentences
Fenglong Xie, Frank K. Soong, Haifeng Li 0001 |
INTERSPEECH | 3 |
| 2015 | An improved ART2 neural network: Resisting pattern drifting through generalized similarity and confidence measures
Haifeng Li 0001, Lin Ma 0003, Yuezhong Song |
Neurocomputing | 1 |
| 2015 | An ICA/HHT Hybrid Approach for Automatic Ocular Artifact CorrectionabstractIn the field of electroencephalogram (EEG) signal processing, the ocular artifact (OA), especially the blink is the most commonly observed, gives the largest disturbance, and can hardly be automatically corrected due to its sudden appearance and wide frequency band influence. In this paper, we focus on a 2-stage independent component analysis/Hilbert Huang transformation (ICA/HHT) hybrid OA correction method which realizes an automatic OA correction as well as a better information conservation to the OA distorted EEG data. In the 1st stage, the ICA is introduced to convert the raw EEG signals into a set of independent signal sources, i.e. a number of independent components (ICs). In the 2nd stage, the HHT is then applied to analyze the ICs in order to emphasize the differences between the OA related ICs and the normal EEG related ICs. On the intrinsic mode functions (IMFs) of each IC extracted by the empirical mode decomposition (EMD), the Hilbert spectrums can clearly indicate where an OA exists and make an automatic OA correction possible. EEG signal samples randomly picked from a variety of neuropsychology tasks were used to evaluate the proposed automatic OA correction method, and the outcomes approved an excellent OA correction performance as well as a better information conservation ability comparing to the well applied methods nowadays. Lin Ma 0003, Haifeng Li 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2014 | Sequence error (SE) minimization training of neural network for voice conversionabstractNeural network (NN) based voice conversion, which employs a nonlinear function to map the features from a source to a target speaker, has been shown to outperform GMM-based voice conversion approach [4-7]. However, there are still limitations to be overcome in NN-based voice conversion, e.g. NN is trained on a Frame Error (FE) minimization criterion and the corresponding weights are adjusted to minimize the error squares over the whole source-target, stereo training data set. In this paper, we use the idea of sentence optimization based, minimum generation error (MGE) training in HMM-based TTS synthesis, and modify the FE minimization to Sequence Error (SE) minimization in NN training for voice conversion. The conversion error over a training sentence from a source speaker to a target speaker is minimized via a gradient descent-based, back propagation (BP) procedure. Experimental results show that the speech converted by the NN, which is first trained with frame error minimization and then refined with sequence error minimization, sounds subjectively better than the converted speech by NN trained with frame error minimization only. Scores on both naturalness and similarity to the target speaker are improved. Index Terms: voice conversion, neural network, pre-training, sequence error minimization Fenglong Xie, Yao Qian, Yuchen Fan 0001, Frank K. Soong, Haifeng Li 0001 |
INTERSPEECH | 5 |
| 2013 | Nonlinear Dynamic Analysis of Pathological Voices
Chunying Fang, Haifeng Li 0001, Lin Ma 0003, Xiaopeng Zhang 0004 |
ICIC (2) | 2 |
| 2013 | Active learning for dimensional speech emotion recognition
Wenjing Han, Haifeng Li 0001, Huabin Ruan, Lin Ma 0003, Jiayin Sun, Björn W. Schuller |
INTERSPEECH | 2 |
| 2013 | A microstructure evolution visualization method based on neutrosophic set theory and cellular automaton techniqueabstractThe visualization of complex physical processes is becoming a challenging topic in fields of information visualization, and attracting more and more researchers. Microstructure evolution visualization (MEV) has become an important and unsubstitutable method in the modern material processing engineering domains. A novel MEV approach combining the neutrosophic set theory (NS) and the cellular automaton technique (CA) is developed in order to precisely simulate the invisible, complex and unrepeatable physical process of metal solidification. The NS theory is applied to realize the complex evolution rules among three phases including solid phase, liquid phase and interface phase, while the CA method is used to simulate the dynamic process of dendrite growth. Experiment results of the dendrite growing simulation show strange consistency of the virtual invisible microstructure with that in practical industrial trials and production. Material experts also convince the statistical characteristics of the simulation results and further the inspirational value for the visualization of other complex physical processes. Haifeng Li 0001, Dayi Yang, Yangyang Fu, Hujie Huang, Hongyuan Fang |
VINCI | 1 |
| 2012 | Preserving actual dynamic trend of emotion in dimensional speech emotion recognitionabstractIn this paper, we use the concept of dynamic trend of emotion to describe how a human's emotion changes over time, which is believed to be important for understanding one's stance toward current topic in interactions. However, the importance of this concept - to our best knowledge - has not been paid enough attention before in the field of speech emotion recognition (SER). Inspired by this, this paper aims to evoke researchers' attention on this concept and makes a primary effort on the research of predicting correct dynamic trend of emotion in the process of SER. Specifically, we propose a novel algorithm named Order Preserving Network (OPNet) to this end. First, as the key issue for OPNet construction, we propose employing a probabilistic method to define an emotion trend-sensitive loss function. Then, a nonlinear neural network is trained using the gradient descent as optimization algorithm to minimize the constructed loss function. We validated the prediction performance of OPNet on the VAM corpus, by mean linear error as well as a rank correlation coefficient γ as measures. Comparing to k-Nearest Neighbor and support vector regression, the proposed OPNet performs better on the preservation of actual dynamic trend of emotion. Wenjing Han, Haifeng Li 0001, Florian Eyben, Lin Ma 0003, Jiayin Sun, Björn W. Schuller |
ICMI | 2 |
| 2012 | An adaptive unsupervised clustering of pronunciation errors for automatic pronunciation error detection
Haifeng Li 0001, Lin Ma 0003 |
ICPR | 2 |
| 2012 | Automatic Pronunciation Error Detection Based on Extended Pronunciation Space Using the Unsupervised Clustering of Pronunciation Errors
Haifeng Li 0001 |
INTERSPEECH | 2 |