VLDB 2026 Research / reviewers in the wild / expert
Xugang Lu
dblp:00/870
· DBLP profile ↗
100ranked-venue papers
30as first author
29since 2021 · last 2026
0000-0001-7075-448XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 81 · 25 first-author · 21 since 2021Artificial intelligence and machine learning · 60 · 21 first-author · 14 since 2021Security and privacy · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Domain Adaptation for Speaker Verification Using Optimal Transport With Pseudo LabelabstractDomain gap often degrades the performance of speaker verification (SV) systems when the statistical distributions of training data and real-world test speech are mismatched. Channel variation is a primary factor causing this gap, including bandwidth changes, background noise and encoding, etc. Although various domain adaptation algorithms could be applied to handle this domain gap problem, most algorithms could not take the complex distribution structure in domain alignment with discriminative learning. In this paper, we propose a novel unsupervised domain adaptation method for speaker verification, i.e., Joint Partial Optimal Transport with Pseudo Label (JPOT-PL), to alleviate the domain mismatch problem. Leveraging the geometric-aware distance metric of optimal transport in distribution alignment and speaker consistency in speech distribution, we further design a pseudo label-based discriminative learning where the pseudo label can be regarded as a new type of speaker label derived from the optimal coupling. With the JPOT-PL, we carry out experiments on the SV channel and lingual domain adaptation with VoxCeleb, LibriSpeech, CNCeleb, and AISHELL-2. Experiments show our method reduces EER by up to 30% compared with several state-of-the-art domain adaptation algorithms. Jianguo Wei, Wenhuan Lu, Lei Li 0050, Xugang Lu |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | Continual Unsupervised Domain Adaptation for Audio Deepfake DetectionabstractAudio deepfake detection (ADD) aims to verify the authenticity of audio. However, its performance declines sharply when facing significant domain discrepancies caused by unknown datasets. Unsupervised domain adaptation (UDA) has been applied to mitigate domain mismatch. However, as generative models evolve, existing UDA methods struggle with catastrophic forgetting when facing continuously emerging spoofing methods. To address this challenge, we introduce continual UDA for ADD, which involves sequentially training across multiple target domains with continual learning. We propose a causality-distillation-based continual domain adversarial training framework for continual UDA, called CD-DAT. Specifically, we employ the domain adversarial training (DAT) framework to learn both spoofing-discriminative and domain-invariant deep features. In addition, we design a continual learning algorithm utilizing causality distillation to capture the mapping between utterances and classes, effectively mitigating forgetting and maintaining generalization. Experiments demonstrated that CD-DAT improved detection performance across all domains, confirming its memory stability and learning plasticity. Xiaohuan Chen, Wenhuan Lu, Ruiteng Zhang, Junhai Xu, Xugang Lu, Lin Zhang 0054, Jianguo Wei |
ICASSP | 5 |
| 2025 | Cross-modal Knowledge Transfer Learning as Graph Matching Based on Optimal Transport for ASR
Xugang Lu, Yu Tsao 0001, Hisashi Kawai |
INTERSPEECH | 1 |
| 2025 | Self-distillation-based domain exploration for source speaker verification under spoofed speech from unknown voice conversion
Xinlei Ma, Ruiteng Zhang, Jianguo Wei, Xugang Lu, Junhai Xu, Lin Zhang 0054, Wenhuan Lu |
Speech Commun. | 4 |
| 2025 | SHDA: Sinkhorn Domain Attention for Cross-Domain Audio Anti-SpoofingabstractAudio anti-spoofing algorithms struggle with fake samples from unseen spoofing techniques, even when trained with diverse data sets or data augmentation strategies. Unsupervised domain adaptation (UDA) algorithms have the potential to mitigate this challenge. Typically, UDA assumes that the source and target domains are distinct distributions with clear boundaries and seeks to align model representations between them. However, in anti-spoofing, various spoofing algorithms could cause the distributions of the generated samples to overlap, resulting in unclear domain boundaries. This hinders UDA algorithms from effectively measuring and aligning domain discrepancies. Moreover, forcibly aligning samples with significant discrepancies could diminish the model’s discriminative capability. To solve this problem, we propose a domain attention algorithm with optimal transport (OT), termed Sinkhorn Domain Attention (SHDA). Unlike traditional attention mechanisms, SHDA identifies the optimal transfer plan by analyzing the global probability differences among cross-domain samples. Specifically, we first extract audio representations from various domains to compute the overall cost matrix between the source and target domains. Next, we employ Sinkhorn’s iteration to calculate the OT coupling matrix, where cross-domain samples with minor differences receive higher transfer weights, while those with substantial differences receive lower weights. Finally, we use the coupling and cost matrices to compute the adaptation loss, effectively transferring the anti-spoofing model from multiple sources to the target domain. We conducted eight cross-domain experiments using eleven well-known anti-spoofing corpora. The results indicate that our label-free SHDA surpassed the state-of-the-art model by 40%. Ruiteng Zhang, Jianguo Wei, Xugang Lu, Lin Zhang 0054, Di Jin 0001, Junhai Xu, Wenhuan Lu |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | MECG-E: Mamba-based ECG Enhancer for Baseline Wander RemovalabstractElectrocardiogram (ECG) is an important non-invasive method for diagnosing cardiovascular disease. However, ECG signals are susceptible to noise contamination, such as electrical interference or signal wandering, which reduces diagnostic accuracy. Various ECG denoising methods have been proposed, but most existing methods yield suboptimal performance under very noisy conditions or require several steps during inference, leading to latency during online processing. In this paper, we propose a novel ECG denoising model, namely Mamba-based ECG Enhancer (MECG-E), which leverages the Mamba architecture known for its fast inference and outstanding nonlinear mapping capabilities. Experimental results indicate that MECG-E surpasses several well-known existing models across multiple metrics under different noise conditions. Additionally, MECG-E requires less inference time than state-of-the-art diffusion-based ECG denoisers, demonstrating the model’s functionality and efficiency. Kuo-Hsuan Hung, Kuan-Chen Wang, Kai-Chun Liu, Wei-Lun Chen, Xugang Lu, Yu Tsao 0001, Chii-Wann Lin |
IEEE Big Data | 5 |
| 2024 | Evaluation of an Improved Ultrasonic Imaging Helmet for Observing Articulatory DataabstractUltrasonic imaging is one of the most popular methods for tracking tongue motion. Imaging plane shift and contact variation are crucial factors affecting the consistency of the obtained ultrasonic images. To solve this issue, researchers proposed many different helmets. In this study, we propose an evaluation framework to quantitatively assess the helmet’s imaging plane shift and contact variation. The framework is applied to one helmet we designed. Compared to the baseline, the imaging plane shift using our helmet is less than 0.1°, and the contact variation is reduced by 68.5%, decreasing the mean difference of the extracted contours by 50.3% and increasing the contrast and sharpness of the obtained images by 9.0% and 28.4%. The proposed framework provides an approach for quantitatively proofing helmet structure with data accuracy and consistency. Jianguo Wei, Qiang Fang 0003, Xugang Lu |
ICASSP | 4 |
| 2024 | Hierarchical Cross-Modality Knowledge Transfer with Sinkhorn Attention for CTC-Based ASRabstractDue to the modality discrepancy between textual and acoustic modeling, efficiently transferring linguistic knowledge from a pretrained language model (PLM) to acoustic encoding for automatic speech recognition (ASR) still remains a challenging task. In this study, we propose a cross-modality knowledge transfer (CMKT) learning framework in a temporal connectionist temporal classification (CTC) based ASR system where hierarchical acoustic alignments with the linguistic representation are applied. Additionally, we propose the use of Sinkhorn attention in cross-modality alignment process, where the transformer attention is a special case of this Sinkhorn attention process. The CMKT learning is supposed to compel the acoustic encoder to encode rich linguistic knowledge for ASR. On the AISHELL-1 dataset, with CTC greedy decoding for inference (without using any language model), we achieved state-of-the-art performance with 3.64% and 3.94% character error rates (CERs) for the development and test sets, which corresponding to relative improvements of 34.18% and 34.88% compared to the baseline CTC-ASR system, respectively. Xugang Lu, Yu Tsao 0001, Hisashi Kawai |
ICASSP | 1 |
| 2024 | Self-Supervised Domain Exploration with an Optimal Transport Regularization for Open Set Cross-Domain Speech Emotion RecognitionabstractIn the tasks of domain adaptation (DA) for speech emotion recognition (SER), self-supervised learning (SSL) algorithms could effectively explore domain and structural information from target domain samples, thereby mitigating domain discrepancies. However, in a general setting, when the target domain contains emotions that are never observed in the source domain, namely in open-set DA, existing SSL-based DA methods cannot maintain the robustness because of the interference of the extra unknown classes. To address this challenge, we propose the self-supervised domain exploration with an optimal transport (OT) regularization (SDEOTR) algorithm. First, we integrate the SSL algorithm into the SER model to mitigate the domain differences. Further, we categorize target domain samples into known and unknown groups based on the network’s prediction confidence. Finally, we employ OT to maximize the global probability distance between the two groups, aiming to decrease the impact of unknown emotions on the SER model. Cross-domain SER experimental results showed that our label-free SDEOTR significantly improved the performance of existing adaptive SER algorithms in open-set scenarios. Ruiteng Zhang, Jianguo Wei, Xugang Lu, Wenhuan Lu, Di Jin 0001, Junhai Xu |
ICASSP | 3 |
| 2024 | Temporal Order Preserved Optimal Transport-Based Cross-Modal Knowledge Transfer Learning for ASRabstractTransferring linguistic knowledge from a pretrained language model (PLM) to an acoustic model has been shown to greatly improve the performance of automatic speech recognition (ASR). However, due to the heterogeneous feature distributions in cross-modalities, designing an effective model for feature alignment and knowledge transfer between linguistic and acoustic sequences remains a challenging task. Optimal transport (OT), which efficiently measures probability distribution discrepancies, holds great potential for aligning and transferring knowledge between acoustic and linguistic modalities. Nonetheless, the original OT treats acoustic and linguistic feature sequences as two unordered sets in alignment and neglects temporal order information during OT coupling estimation. Consequently, a time-consuming pretraining stage is required to learn a good alignment between the acoustic and linguistic representations. In this paper, we propose a Temporal Order Preserved OT (TOT)-based Cross-modal Alignment and Knowledge Transfer (CAKT) (TOT-CAKT) for ASR. In the TOT-CAKT, local neighboring frames of acoustic sequences are smoothly mapped to neighboring regions of linguistic sequences, preserving their temporal order relationship in feature alignment and matching. With the TOT-CAKT model framework, we conduct Mandarin ASR experiments with a pretrained Chinese PLM for linguistic knowledge transfer. Our results demonstrate that the proposed TOT-CAKT significantly improves ASR performance compared to several state-of-the-art models employing linguistic knowledge transfer, and addresses the weaknesses of the original OT-based method in sequential feature alignment for ASR. Xugang Lu, Yu Tsao 0001, Hisashi Kawai |
SLT | 1 |
| 2024 | Distillation-Based Feature Extraction Algorithm For Source Speaker VerificationabstractAutomatic speaker verification (ASV) systems face significant challenges when exposed to spoofing attacks, necessitating robust countermeasures. In this work, we focus on the source speaker verification (SSV) task, which aims to identify the source speaker hidden in spoofed speech generated by voice conversion (VC) systems. We propose a distillation-based feature extraction algorithm to enhance the model’s ability to verify source speakers. Our method employs a pretraining ASV model as a teacher network and the SSV model as a student network, using bona fide speech to guide the learning process. However, the improvements were marginal, particularly on the development set, indicating the complexity and resource demands of fine-tuning the distillation parameters. Our findings underscore the inherent difficulties in SSV and highlight the need for further research to develop more effective solutions. Besides, our submission won fourth place in the 2024 Source Speaker Tracking Challenge. Xinlei Ma, Wenhuan Lu, Ruiteng Zhang, Junhai Xu, Xugang Lu, Jianguo Wei |
SLT | 5 |
| 2024 | Unsupervised Adaptive Speaker Recognition by Coupling-Regularized Optimal TransportabstractCross-domain speaker recognition (SR) can be improved by unsupervised domain adaptation (UDA) algorithms. UDA algorithms often reduce domain mismatch at the cost of decreasing the discrimination of speaker features. In contrast, optimal transport (OT) has the potential to achieve domain alignment while preserving the speaker discrimination capability in UDA applications; however, naively applying OT to measure global probability distribution discrepancies between the source and target domains may induce negative transports where samples belonging to different speakers are coupled in transportation. These negative transports reduce the SR model's discriminative power, degrading the SR performance. This paper proposes a coupling-regularized optimal transport (CROT) algorithm for cross-domain SR to reduce the negative transport during UDA. In the proposed CROT, two consecutive processing modules regularize the coupling paths for the OT solution: a progressive inter-speaker constraint (PISC) module and a coupling-smoothed regularization (CSR) module. The PISC, designed as a pseudo-label memory bank with curriculum learning, is first applied to select valid samples to guarantee that coupling samples are from the same speaker. The CSR, designed to control the information entropy of the coupling paths further, reduces the effect of negative transport in UDA. To evaluate the effectiveness of the proposed algorithm, cross-domain SR experiments were conducted under different target domains, speaker encoders, corpora, and acoustic features. Experimental results showed that CROT achieved a 50% relative reduction in equal error rates compared to conventional OT-based UDAs, outperforming the state-of-the-art UDAs. Ruiteng Zhang, Jianguo Wei, Xugang Lu, Wenhuan Lu, Di Jin 0001, Lin Zhang 0054, Junhai Xu |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Cross-Modal Alignment With Optimal Transport For CTC-Based ASRabstractTemporal connectionist temporal classification (CTC)-based automatic speech recognition (ASR) is one of the most successful end to end (E2E) ASR frameworks. However, due to the token independence assumption in decoding, an external language model (LM) is required which destroys its fast parallel decoding property. Several studies have been proposed to transfer linguistic knowledge from a pretrained LM (PLM) to the CTC based ASR. Since the PLM is built from text while the acoustic model is trained with speech, a cross-modal alignment is required in order to transfer the context dependent linguistic knowledge from the PLM to acoustic encoding. In this study, we propose a novel cross-modal alignment algorithm based on optimal transport (OT). In the alignment process, a transport coupling matrix is obtained using OT, which is then utilized to transform a latent acoustic representation for matching the context-dependent linguistic features encoded by the PLM. Based on the alignment, the latent acoustic feature is forced to encode context dependent linguistic information. We integrate this latent acoustic feature to build conformer encoder-based CTC ASR system. On the AISHELL-1 data corpus, our system achieved 3.96 % and 4.27 % character error rate (CER) for dev and test sets, respectively, which corresponds to relative improvements of 28.39 % and 29.42% compared to the baseline conformer CTC ASR system without cross-modal knowledge transfer. Xugang Lu, Yu Tsao 0001, Hisashi Kawai |
ASRU | 1 |
| 2023 | Optimal Transport with a Diversified Memory Bank for Cross-Domain Speaker VerificationabstractOptimal transport (OT) can be applied to cross-domain adaptation in speaker verification (SV) by converting speakers' probability distributions from source to target domains. However, in scenarios involving over-massive categories (speakers) or difficult samples in discrimination, OT often has difficulty computing effective transports. To address this challenge, we propose an OT-based unsupervised domain adaptation (UDA) framework for SV, OT with a diversified memory bank, called DMB-OT, which ensures the accuracy of transfers by two strategies: (1) It regularizes the solution space of OT, which attempts to plan transformations between audio samples from the same speaker with high confidence; (2) it integrates a dynamic curriculum learning algorithm, preventing OT from calculating transport couplings based on hard-discriminative samples in the early stage of UDA. Experiments under different target domains showed that our unsupervised DMB-OT could significantly improve the performance of OT-based UDA and could even match the performance of the supervised PLDA-based adaptation. Ruiteng Zhang, Jianguo Wei, Xugang Lu, Wenhuan Lu, Di Jin 0001, Lin Zhang 0054, Junhai Xu |
ICASSP | 3 |
| 2023 | Multi-Level Knowledge Distillation for Speech Emotion Recognition in Noisy Conditions
Yang Liu 0262, Haoqin Sun, Qingyue Wang, Zhen Zhao 0006, Xugang Lu, Longbiao Wang |
INTERSPEECH | 6 |
| 2023 | SOT: Self-supervised Learning-Assisted Optimal Transport for Unsupervised Adaptive Speech Emotion Recognition
Ruiteng Zhang, Jianguo Wei, Xugang Lu, Junhai Xu, Di Jin 0001, Jianhua Tao 0001 |
INTERSPEECH | 3 |
| 2023 | TMS: Temporal multi-scale in time-delay neural network for speaker verification
Ruiteng Zhang, Jianguo Wei, Xugang Lu, Wenhuan Lu, Di Jin 0001, Lin Zhang 0054, Junhai Xu, Jianwu Dang 0001 |
Appl. Intell. | 3 |
| 2023 | Self-supervised learning based domain regularization for mask-wearing speaker verification
Ruiteng Zhang, Jianguo Wei, Xugang Lu, Wenhuan Lu, Di Jin 0001, Lin Zhang 0054, Yantao Ji, Junhai Xu |
Speech Commun. | 3 |
| 2022 | CS-REP: Making Speaker Verification Networks Embracing Re-ParameterizationabstractAutomatic speaker verification (ASV) systems, which determine whether two speeches are from the same speaker, mainly focus on verification accuracy while ignoring inference speed. However, in real applications, both inference speed and verification accuracy are essential. This study proposes cross-sequential re-parameterization (CS-Rep), a novel topology re-parameterization strategy for multi-type networks, to increase the inference speed and verification accuracy of models. CS-Rep solves the problem that existing re-parameterization methods are not suitable for typical ASV backbones. When a model applies CS-Rep, the training-period network utilizes a multi-branch topology to capture speaker information, whereas the inference-period model converts to a time-delay neural network (TDNN)-like plain backbone with stacked TDNN layers to achieve the fast inference speed. Based on CS-Rep, an improved TDNN with friendly test and deployment called Rep-TDNN is proposed. Compared with the state-of-the-art model ECAPA-TDNN, Rep-TDNN increases the actual inference speed by about 50% and reduces the EER by 10%. The code and trained models are available at https://github.com/zrtlemontree/CS-Rep. Ruiteng Zhang, Jianguo Wei, Wenhuan Lu, Lin Zhang 0054, Yantao Ji, Junhai Xu, Xugang Lu |
ICASSP | 7 |
| 2022 | Perceptual Contrast Stretching on Target Feature for Speech EnhancementabstractSpeech enhancement (SE) performance has improved considerably owing to the use of deep learning models as a base function.Herein, we propose a perceptual contrast stretching (PCS) approach to further improve SE performance.The PCS is derived based on the critical band importance function and is applied to modify the targets of the SE model.Specifically, the contrast of target features is stretched based on perceptual importance, thereby improving the overall SE performance.Compared with post-processing-based implementations, incorporating PCS into the training phase preserves performance and reduces online computation.Notably, PCS can be combined with different SE model architectures and training criteria.Furthermore, PCS does not affect the causality or convergence of SE model training.Experimental results on the VoiceBank-DEMAND dataset show that the proposed method can achieve state-of-the-art performance on both causal (PESQ score = 3.07) and noncausal (PESQ score = 3.35) SE tasks. Rong Chao, Szu-Wei Fu, Xugang Lu, Yu Tsao 0001 |
INTERSPEECH | 4 |
| 2022 | Data Augmentation Using McAdams-Coefficient-Based Speaker Anonymization for Fake Audio DetectionabstractFake audio detection (FAD) is a technique to distinguish synthetic speech from natural speech.In most FAD systems, removing irrelevant features from acoustic speech while keeping only robust discriminative features is essential.Intuitively, speaker information entangled in acoustic speech should be suppressed for the FAD task.Particularly in a deep neural network (DNN)-based FAD system, the learning system may learn speaker information from a training dataset and cannot generalize well on a testing dataset.In this paper, we propose to use the speaker anonymization (SA) technique to suppress speaker information from acoustic speech before inputting it into a DNN-based FAD system.We adopted the McAdamscoefficient-based SA (MC-SA) algorithm, and this is expected that the entangled speaker information will not be involved in the DNN-based FAD learning.Based on this idea, we implemented a light convolutional neural network bidirectional long short-term memory (LCNN-BLSTM)-based FAD system and conducted experiments on the Audio Deep Synthesis Detection Challenge (ADD2022) datasets.The results showed that removing the speaker information from acoustic speech improved the relative performance in the first track of ADD2022 by 17.66%. Kai Li 0018, Sheng Li 0010, Xugang Lu, Masato Akagi, Meng Liu 0017, Lin Zhang 0054, Chang Zeng, Longbiao Wang, Jianwu Dang 0001, Masashi Unoki |
INTERSPEECH | 3 |
| 2022 | Transducer-based language embedding for spoken language identificationabstractThe acoustic and linguistic features are important cues for the spoken language identification (LID) task.Recent advanced LID systems mainly use acoustic features that lack the usage of explicit linguistic feature encoding.In this paper, we propose a novel transducer-based language embedding approach for LID tasks by integrating an RNN transducer model into a language embedding framework.Benefiting from the advantages of the RNN transducer's linguistic representation capability, the proposed method can exploit both phonetically-aware acoustic features and explicit linguistic features for LID tasks.Experiments were carried out on the large-scale multilingual LibriSpeech and VoxLingua107 datasets.Experimental results showed the proposed method significantly improves the performance on LID tasks with 12% to 59% and 16% to 24% relative improvement on in-domain and cross-domain datasets, respectively. Xugang Lu, Hisashi Kawai |
INTERSPEECH | 2 |
| 2022 | Pronunciation-Aware Unique Character Encoding for RNN Transducer-Based Mandarin Speech RecognitionabstractFor Mandarin end-to-end (E2E) automatic speech recognition (ASR) tasks, compared to character-based modeling units, pronunciation-based modeling units could improve the sharing of modeling units in model training but meet homophone problems. In this study, we propose to use a novel pronunciation-aware unique character encoding for building E2E RNN-T-based Mandarin ASR systems. The proposed encoding is a combination of pronunciation-base syllable and character index (CI). By introducing the CI, the RNN-T model can overcome the homophone problem while utilizing the pronunciation information for extracting modeling units. With the proposed encoding, the model outputs can be converted into the final recognition result through a one-to-one mapping. We conducted experiments on Aishell and MagicData datasets, and the experimental results showed the effectiveness of the proposed method. Xugang Lu, Hisashi Kawai |
SLT | 2 |
| 2021 | Unsupervised Neural Adaptation Model Based on Optimal Transport for Spoken Language IdentificationabstractDue to the mismatch of statistical distributions of acoustic speech between training and testing sets, the performance of spoken language identification (SLID) could be drastically degraded. In this paper, we propose an unsupervised neural adaptation model to deal with the distribution mismatch problem for SLID. In our model, we explicitly formulate the adaptation as to reduce the distribution discrepancy on both feature and classifier for training and testing data sets. Moreover, inspired by the strong power of the optimal transport (OT) to measure distribution discrepancy, a Wasserstein distance metric is designed in the adaptation loss. By minimizing the classification loss on the training data set with the adaptation loss on both training and testing data sets, the statistical distribution difference between training and testing domains is reduced. We carried out SLID experiments on the oriental language recognition (OLR) challenge data corpus where the training and testing data sets were collected from different conditions. Our results showed that significant improvements were achieved on the cross domain test tasks. Xugang Lu, Yu Tsao 0001, Hisashi Kawai |
ICASSP | 1 |
| 2021 | MetricGAN+: An Improved Version of MetricGAN for Speech EnhancementabstractThe discrepancy between the cost function used for training a speech enhancement model and human auditory perception usually makes the quality of enhanced speech unsatisfactory.Objective evaluation metrics which consider human perception can hence serve as a bridge to reduce the gap.Our previously proposed MetricGAN was designed to optimize objective metrics by connecting the metric with a discriminator.Because only the scores of the target evaluation functions are needed during training, the metrics can even be non-differentiable.In this study, we propose a MetricGAN+ in which three training techniques incorporating domainknowledge of speech processing are proposed.With these techniques, experimental results on the VoiceBank-DEMAND dataset show that MetricGAN+ can increase PESQ score by 0.3 compared to the previous MetricGAN and achieve stateof-the-art results (PESQ score = 3.15). Szu-Wei Fu, Tsun-An Hsieh, Peter Plantinga, Mirco Ravanelli, Xugang Lu, Yu Tsao 0001 |
Interspeech | 6 |
| 2021 | Improving Perceptual Quality by Phone-Fortified Perceptual Loss Using Wasserstein Distance for Speech EnhancementabstractSpeech enhancement (SE) aims to improve speech quality and intelligibility, which are both related to a smooth transition in speech segments that may carry linguistic information, e.g.phones and syllables.In this study, we propose a novel phonefortified perceptual loss (PFPL) that takes phonetic information into account for training SE models.To effectively incorporate the phonetic information, the PFPL is computed based on latent representations of the wav2vec model, a powerful selfsupervised encoder that renders rich phonetic information.To more accurately measure the distribution distances of the latent representations, the PFPL adopts the Wasserstein distance as the distance measure.Our experimental results first reveal that the PFPL is more correlated with the perceptual evaluation metrics, as compared to signal-level losses.Moreover, the results showed that the PFPL can enable a deep complex U-Net SE model to achieve highly competitive performance in terms of standardized quality and intelligibility evaluations on the Voice Bank-DEMAND dataset. Tsun-An Hsieh, Szu-Wei Fu, Xugang Lu, Yu Tsao 0001 |
Interspeech | 4 |
| 2021 | EMA2S: An End-to-End Multimodal Articulatory-to-Speech SystemabstractSynthesized speech from articulatory movements can have real-world use for patients with vocal cord disorders, situations requiring silent speech, or in high-noise environments. In this work, we present EMA2S, an end-to-end multimodal articulatory-to-speech system that directly converts articulatory movements to speech signals. We use a neural-network-based vocoder combined with multimodal joint-training, incorporating spectrogram, mel-spectrogram, and deep features. The experimental results confirm that the multimodal approach of EMA2S outperforms the baseline system in terms of both objective evaluation and subjective evaluation metrics. Moreover, results demonstrate that joint mel-spectrogram and deep feature loss training can effectively improve system performance. Yuwen Chen 0006, Kuo-Hsuan Hung, Shang-Yi Chuang, Jonathan Sherman, Wen-Chin Huang, Xugang Lu, Yu Tsao 0001 |
ISCAS | 6 |
| 2021 | Unsupervised Noise Adaptive Speech Enhancement by Discriminator-Constrained Optimal TransportabstractThis paper presents a novel discriminator-constrained optimal transport network (DOTN) that performs unsupervised domain adaptation for speech enhancement (SE), which is an essential regression task in speech processing. The DOTN aims to estimate clean references of noisy speech in a target domain, by exploiting the knowledge available from the source domain. The domain shift between training and testing data has been reported to be an obstacle to learning problems in diverse fields. Although rich literature exists on unsupervised domain adaptation for classification, the methods proposed, especially in regressions, remain scarce and often depend on additional information regarding the input data. The proposed DOTN approach tactically fuses the optimal transport (OT) theory from mathematical analysis with generative adversarial frameworks, to help evaluate continuous labels in the target domain. The experimental results on two SE tasks demonstrate that by extending the classical OT formulation, our proposed DOTN outperforms previous adversarial domain adaptation frameworks in a purely unsupervised manner. Hsin-Yi Lin, Huan-Hsin Tseng, Xugang Lu, Yu Tsao 0001 |
NeurIPS | 3 |
| 2021 | Coupling a Generative Model With a Discriminative Learning Framework for Speaker VerificationabstractThe task of speaker verification (SV) is to decide whether an utterance is spoken by a target or an imposter speaker. In most studies of SV, a log-likelihood ratio (LLR) score is estimated based on a generative probability model on speaker features, and compared with a threshold for making a decision. However, the generative model usually focuses on individual feature distributions, does not have the discriminative feature selection ability, and is easy to be distracted by nuisance features. The SV, as a hypothesis test, could be formulated as a binary discrimination task where neural network based discriminative learning could be applied. In discriminative learning, the nuisance features could be removed with the help of label supervision. However, discriminative learning pays more attention to classification boundaries, and is prone to overfitting to a training set which may result in bad generalization on a test set. In this paper, we propose a hybrid learning framework, i.e., coupling a joint Bayesian (JB) generative model structure and parameters with a neural discriminative learning framework for SV. In the hybrid framework, a two-branch Siamese neural network is built with dense layers that are coupled with factorized affine transforms as used in the JB model. The LLR score estimation in the JB model is formulated according to the distance metric in the discriminative learning framework. By initializing the two-branch neural network with the generatively learned model parameters of the JB model, we further train the model parameters with the pairwise samples as a binary discrimination task. Moreover, a direct evaluation metric (DEM) in SV based on minimum empirical Bayes risk (EBR) is designed and integrated as an objective function in the discriminative learning. We carried out SV experiments on Speakers in the wild (SITW) and Voxceleb. Experimental results showed that our proposed model improved the performance with a large margin compared with state of the art models for SV. Xugang Lu, Yu Tsao 0001, Hisashi Kawai |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2020 | Robust Unsupervised Neural Machine Translation with Adversarial Denoising TrainingabstractUnsupervised neural machine translation (UNMT) has recently attracted great interest in the machine translation community.The main advantage of the UNMT lies in its easy collection of required large training text sentences while with only a slightly worse performance than supervised neural machine translation which requires expensive annotated translation pairs on some translation tasks.In most studies, the UMNT is trained with clean data without considering its robustness to the noisy data.However, in real-world scenarios, there usually exists noise in the collected input sentences which degrades the performance of the translation system since the UNMT is sensitive to the small perturbations of the input sentences.In this paper, we first time explicitly take the noisy data into consideration to improve the robustness of the UNMT based systems.First of all, we clearly defined two types of noises in training sentences, i.e., word noise and word order noise, and empirically investigate its effect in the UNMT, then we propose adversarial training methods with denoising process in the UNMT.Experimental results on several language pairs show that our proposed methods substantially improved the robustness of the conventional UNMT systems in noisy scenarios. Haipeng Sun, Rui Wang 0015, Kehai Chen, Xugang Lu, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
COLING | 4 |
| 2020 | Self-Supervised Denoising Autoencoder with Linear Regression Decoder for Speech EnhancementabstractNonlinear spectral mapping-based models based on supervised learning have successfully applied for speech enhancement. However, as supervised learning approaches, a large amount of labelled data (noisy-clean speech pairs) should be provided to train those models. In addition, their performances for unseen noisy conditions are not guaranteed, which is a common weak point of supervised learning approaches. In this study, we proposed an unsupervised learning approach for speech enhancement, i.e., denoising autoencoder with linear regression decoder (DAELD) model for speech enhancement. The DAELD is trained with noisy speech as both input and target output in a self-supervised learning manner. In addition, with properly setting a shrinkage threshold for internal hidden representations, noise could be removed during the reconstruction from the hidden representations via the linear regression decoder. Speech enhancement experiments were carried out to test the proposed model. Results confirmed that the proposed DAELD could achieve comparable and sometimes even better enhancement performance as compared to the conventional supervised speech enhancement approaches, in both seen and unseen noise environments. Moreover, we observe that higher performances tend to achieve by DAELD when the training data cover more diverse noise types and signal-tonoise-ratio (SNR) levels. Ryandhimas E. Zezario, Tassadaq Hussain, Xugang Lu, Hsin-Min Wang, Yu Tsao 0001 |
ICASSP | 3 |
| 2020 | Incorporating Broad Phonetic Information for Speech EnhancementabstractIn noisy conditions, knowing speech contents facilitates listeners to more effectively suppress background noise components and to retrieve pure speech signals.Previous studies have also confirmed the benefits of incorporating phonetic information in a speech enhancement (SE) system to achieve better denoising performance.To obtain the phonetic information, we usually prepare a phoneme-based acoustic model, which is trained using speech waveforms and phoneme labels.Despite performing well in normal noisy conditions, when operating in very noisy conditions, however, the recognized phonemes may be erroneous and thus misguide the SE process.To overcome the limitation, this study proposes to incorporate the broad phonetic class (BPC) information into the SE process.We have investigated three criteria to build the BPC, including two knowledgebased criteria: place and manner of articulatory and one datadriven criterion.Moreover, the recognition accuracies of BPCs are much higher than that of phonemes, thus providing more accurate phonetic information to guide the SE process under very noisy conditions.Experimental results demonstrate that the proposed SE with the BPC information framework can achieve notable performance improvements over the baseline system and an SE system using monophonic information in terms of both speech quality intelligibility on the TIMIT dataset. Yen-Ju Lu, Chien-Feng Liao, Xugang Lu, Jeih-Weih Hung, Yu Tsao 0001 |
INTERSPEECH | 3 |
| 2020 | Investigation of NICT Submission for Short-Duration Speaker Verification Challenge 2020
Xugang Lu, Hisashi Kawai |
INTERSPEECH | 2 |
| 2020 | WaveCRN: An Efficient Convolutional Recurrent Neural Network for End-to-End Speech EnhancementabstractDue to the simple design pipeline, end-to-end (E2E) neural models for speech enhancement (SE) have attracted great interest. In order to improve the performance of the E2E model, the local and sequential properties of speech should be efficiently taken into account when modelling. However, in most current E2E models for SE, these properties are either not fully considered or are too complex to be realized. In this letter, we propose an efficient E2E SE model, termed WaveCRN. Compared with models based on convolutional neural networks (CNN) or long short-term memory (LSTM), WaveCRN uses a CNN module to capture the speech locality features and a stacked simple recurrent units (SRU) module to model the sequential property of the locality features. Different from conventional recurrent neural networks and LSTM, SRU can be efficiently parallelized in calculation, with even fewer model parameters. In order to more effectively suppress noise components in the noisy speech, we derive a novel restricted feature masking approach, which performs enhancement on the feature maps in the hidden layers; this is different from the approaches that apply the estimated ratio mask to the noisy spectral features, which is commonly used in speech separation methods. Experimental results on speech denoising and compressed speech restoration tasks confirm that with the SRU and the restricted feature map, WaveCRN performs comparably to other state-of-the-art approaches with notably reduced model complexity and inference time. Tsun-An Hsieh, Hsin-Min Wang, Xugang Lu, Yu Tsao 0001 |
IEEE Signal Process. Lett. | 3 |
| 2020 | Knowledge Distillation-Based Representation Learning for Short-Utterance Spoken Language IdentificationabstractWith successful applications of deep feature learning algorithms, spoken language identification (LID) on long utterances obtains satisfactory performance. However, the performance on short utterances is drastically degraded even when the LID system is trained using short utterances. The main reason is due to the large variation of the representation on short utterances which results in high model confusion. To narrow the performance gap between long, and short utterances, we proposed a teacher-student representation learning framework based on a knowledge distillation method to improve LID performance on short utterances. In the proposed framework, in addition to training the student model on short utterances with their true labels, the internal representation from the output of a hidden layer of the student model is supervised with the representation corresponding to their longer utterances. By reducing the distance of internal representations between short, and long utterances, the student model can explore robust discriminative representations for short utterances, which is expected to reduce model confusion. We conducted experiments on our in-house LID dataset, and NIST LRE07 dataset, and showed the effectiveness of the proposed methods for short utterance LID tasks. Xugang Lu, Sheng Li 0010, Hisashi Kawai |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Speech Enhancement Based on Denoising Autoencoder With Multi-Branched EncodersabstractDeep learning-based models have greatly advanced the performance of speech enhancement (SE) systems. However, two problems remain unsolved, which are closely related to model generalizability to noisy conditions: (1) mismatched noisy condition during testing, i.e., the performance is generally sub-optimal when models are tested with unseen noise types that are not involved in the training data; (2) local focus on specific noisy conditions, i.e., models trained using multiple types of noises cannot optimally remove a specific noise type even though the noise type has been involved in the training data. These problems are common in real applications. In this article, we propose a novel denoising autoencoder with a multi-branched encoder (termed DAEME) model to deal with these two problems. In the DAEME model, two stages are involved: training and testing. In the training stage, we build multiple component models to form a multi-branched encoder based on a decision tree (DSDT). The DSDT is built based on prior knowledge of speech and noisy conditions (the speaker, environment, and signal factors are considered in this paper), where each component of the multi-branched encoder performs a particular mapping from noisy to clean speech along the branch in the DSDT. Finally, a decoder is trained on top of the multi-branched encoder. In the testing stage, noisy speech is first processed by each component model. The multiple outputs from these models are then integrated into the decoder to determine the final enhanced speech. Experimental results show that DAEME is superior to several baseline models in terms of objective evaluation metrics, automatic speech recognition results, and quality in subjective human listening tests. Ryandhimas E. Zezario, Syu-Siang Wang, Jonathan Sherman, Yi-Yen Hsieh, Xugang Lu, Hsin-Min Wang, Yu Tsao 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2019 | Interactive Learning of Teacher-student Model for Short Utterance Spoken Language IdentificationabstractShort utterance-based spoken language identification (LID) is a challenging task due to the large variation of its feature representation. Improving feature representation of short utterances using a teacher-student method has been shown its effectiveness for LID tasks. However, conventional teacher-student methods use fixed pre-trained teacher models, that makes it difficult to optimize student models. In this paper, rather than using a fixed pre-trained teacher model, we investigate an interactive teacher-student learning by adjusting the teacher model with reference to the performance of the student model when the student model is stuck in a local minimum. Experiments on a 10-language LID task were carried out to test the algorithm. Our results showed its effectiveness of the proposed algorithm on short utterance LID tasks. Xugang Lu, Sheng Li 0010, Hisashi Kawai |
ICASSP | 2 |
| 2019 | End-to-End Articulatory Attribute Modeling for Low-Resource Multilingual Speech Recognition
Sheng Li 0010, Chenchen Ding, Xugang Lu, Tatsuya Kawahara, Hisashi Kawai |
INTERSPEECH | 3 |
| 2019 | Investigating Radical-Based End-to-End Speech Recognition Systems for Chinese Dialects and Japanese
Sheng Li 0010, Xugang Lu, Chenchen Ding, Tatsuya Kawahara, Hisashi Kawai |
INTERSPEECH | 2 |
| 2019 | Improving Transformer-Based Speech Recognition Systems with Compressed Structure and Speech Attributes Augmentation
Sheng Li 0010, Raj Dabre, Xugang Lu, Tatsuya Kawahara, Hisashi Kawai |
INTERSPEECH | 3 |
| 2019 | Incorporating Symbolic Sequential Modeling for Speech EnhancementabstractIn a noisy environment, a lossy speech signal can be automatically restored by a listener if he/she knows the language well.That is, with the built-in knowledge of a "language model", a listener may effectively suppress noise interference and retrieve the target speech signals.Accordingly, we argue that familiarity with the underlying linguistic content of spoken utterances benefits speech enhancement (SE) in noisy environments.In this study, in addition to the conventional modeling for learning the acoustic noisy-clean speech mapping, an abstract symbolic sequential modeling is incorporated into the SE framework.This symbolic sequential modeling can be regarded as a "linguistic constraint" in learning the acoustic noisy-clean speech mapping function.In this study, the symbolic sequences for acoustic signals are obtained as discrete representations with a Vector Quantized Variational Autoencoder algorithm.The obtained symbols are able to capture high-level phoneme-like content from speech signals.The experimental results demonstrate that the proposed framework can obtain notable performance improvement in terms of perceptual evaluation of speech quality (PESQ) and short-time objective intelligibility (STOI) on the TIMIT dataset. Chien-Feng Liao, Yu Tsao 0001, Xugang Lu, Hisashi Kawai |
INTERSPEECH | 3 |
| 2019 | Class-Wise Centroid Distance Metric Learning for Acoustic Event Detection
Xugang Lu, Sheng Li 0010, Yu Tsao 0001, Hisashi Kawai |
INTERSPEECH | 1 |
| 2019 | Specialized Speech Enhancement Model Selection Based on Learned Non-Intrusive Quality Assessment Metric
Ryandhimas E. Zezario, Szu-Wei Fu, Xugang Lu, Hsin-Min Wang, Yu Tsao 0001 |
INTERSPEECH | 3 |
| 2018 | Speech Dereverberation Based on Integrated Deep and Ensemble Learning AlgorithmabstractReverberation, which is generally caused by sound reflections from walls, ceilings, and floors, can result in severe performance degradation of acoustic applications. Due to a complicated combination of attenuation and time-delay effects, the reverberation property is difficult to characterize, and it remains a challenging task to effectively retrieve the anechoic speech signals from reverberation ones. In the present study, we proposed a novel integrated deep and ensemble learning algorithm (IDEA) for speech dereverberation. The IDEA consists of offline and online phases. In the offline phase, we train multiple dereverberation models, each aiming to precisely dereverb speech signals in a particular acoustic environment; then a unified fusion function is estimated that aims to integrate the information of multiple dereverberation models. In the online phase, an input utterance is first processed by each of the dereverberation models. The outputs of all models are integrated accordingly to generate the final anechoic signal. We evaluated the IDEA on designed acoustic environments, including both matched and mismatched conditions of the training and testing data. Experimental results confirm that the proposed IDEA outperforms single deep-neural-network-based dereverberation model with the same model architecture and training data. Wei-Jen Lee, Syu-Siang Wang, Fei Chen 0011, Xugang Lu, Shao-Yi Chien, Yu Tsao 0001 |
ICASSP | 4 |
| 2018 | Improving CTC-based Acoustic Model with Very Deep Residual Time-delay Neural Networks
Sheng Li 0010, Xugang Lu, Ryoichi Takashima, Tatsuya Kawahara, Hisashi Kawai |
INTERSPEECH | 2 |
| 2018 | Temporal Attentive Pooling for Acoustic Event Detection
Xugang Lu, Sheng Li 0010, Yu Tsao 0001, Hisashi Kawai |
INTERSPEECH | 1 |
| 2018 | Feature Representation of Short Utterances Based on Knowledge Distillation for Spoken Language Identification
Xugang Lu, Sheng Li 0010, Hisashi Kawai |
INTERSPEECH | 2 |
| 2018 | Improving Very Deep Time-Delay Neural Network With Vertical-Attention For Effectively Training CTC-Based ASR SystemsabstractThe very deep neural network has recently been proposed for speech recognition and achieves significant performance. It has excellent potential for integration with end-to-end (E2E) training. Connectionist temporal classification (CTC) has shown great potential in E2E acoustic modeling. In this study, we investigate deep architectures and techniques which are suitable for CTC-based acoustic modeling. We propose a very deep residual time-delay CTC neural network (VResTD-CTC). How to select a suitable deep architecture optimized with the CTC objective function is crucial for obtaining the state of the art performance. Excellent performances can be obtained by selecting deep architecture for non-E2E ASR systems modeling with tied-triphone states. However, these optimized structures do not guarantee to achieve better or comparable performances on E2E (e.g., CTC-based) systems modeling with dynamic acoustic units. For solving this problem and further leveraging the system performance, we introduce the vertical-attention mechanism to reweight the residual blocks at each time step. Speech recognition experiments show our proposed model significantly outperforms the DNN and LSTM-based (both bidirectional and unidirectional) CTC baseline models. Sheng Li 0010, Xugang Lu, Ryoichi Takashima, Tatsuya Kawahara, Hisashi Kawai |
SLT | 2 |
| 2018 | Study of articulators' contribution and compensation during speech by articulatory speech recognition
Jianguo Wei, Jingshu Zhang, Qiang Fang 0003, Wenhuan Lu, Kiyoshi Honda, Xugang Lu |
Multim. Tools Appl. | 7 |
| 2018 | End-to-End Waveform Utterance Enhancement for Direct Evaluation Metrics Optimization by Fully Convolutional Neural NetworksabstractSpeech enhancement model is used to map a noisy speech to a clean speech. In the training stage, an objective function is often adopted to optimize the model parameters. However, in the existing literature, there is an inconsistency between the model optimization criterion and the evaluation criterion for the enhanced speech. For example, in measuring speech intelligibility, most of the evaluation metric is based on a short-time objective intelligibility (STOI) measure, while the frame based mean square error (MSE) between estimated and clean speech is widely used in optimizing the model. Due to the inconsistency, there is no guarantee that the trained model can provide optimal performance in applications. In this study, we propose an end-to-end utterance-based speech enhancement framework using fully convolutional neural networks (FCN) to reduce the gap between the model optimization and the evaluation criterion. Because of the utterance-based optimization, temporal correlation information of long speech segments, or even at the entire utterance level, can be considered to directly optimize perception-based objective functions. As an example, we implemented the proposed FCN enhancement framework to optimize the STOI measure. Experimental results show that the STOI of a test speech processed by the proposed approach is better than conventional MSE-optimized speech due to the consistency between the training and the evaluation targets. Moreover, by integrating the STOI into model optimization, the intelligibility of human subjects and automatic speech recognition system on the enhanced speech is also substantially improved compared to those generated based on the minimum MSE criterion. Szu-Wei Fu, Taowei Wang, Yu Tsao 0001, Xugang Lu, Hisashi Kawai |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2017 | Incremental training and constructing the very deep convolutional residual network acoustic modelsabstractInspired by the successful applications in image recognition, the very deep convolutional residual network (ResNet) based model has been applied in automatic speech recognition (ASR). However, the computational load is heavy for training the ResNet with a large quantity of data. In this paper, we propose an incremental model training framework to accelerate the training process of the ResNet. The incremental model training framework is based on the unequal importance of each layer and connection in the ResNet. The modules with important layers and connections are regarded as a skeleton model, while those left are regarded as an auxiliary model. The total depth of the skeleton model is quite shallow compared to the very deep full network. In our incremental training, the skeleton model is first trained with the full training data set. Other layers and connections belonging to the auxiliary model are gradually attached to the skeleton model and tuned. Our experiments showed that the proposed incremental training obtained comparable performances and faster training speed compared with the model training as a whole without consideration of the different importance of each layer. Sheng Li 0010, Xugang Lu, Ryoichi Takashima, Tatsuya Kawahara, Hisashi Kawai |
ASRU | 2 |
| 2017 | Minimum Bayes risk training of CTC acoustic models in maximum a posteriori based decoding frameworkabstractWhen using connectionist temporal classification (CTC) based acoustic models (AMs) for large vocabulary continuous speech recognition (LVCSR), most previous studies have used a naive interpolation of the CTC-AM score and an additional language model score, although there is no theoretical justification for such an approach. On the other hand, we recently proposed a theoretically more sound decoding framework for CTC-AM called maximum a posteriori (MAP)-based decoding. Although the superiority of the MAP-based decoding framework with CTC-AM has been demonstrated, the effect of additional minimum Bayes risk (MBR) training in the MAP-based decoding framework has not been investigated. In this paper, we report the results of various experiments that examine the effect of MBR training on CTC-AM by comparing two decoding frameworks. Our experiments with English and Japanese LVCSR tasks reveal that the MAP-based decoding framework is superior to the interpolation-based framework, even after the MBR training. In addition, by using about 600 h of training data, we show that the size of the training dataset is a critical factor in achieving good results under CTC-AM. Naoyuki Kanda, Xugang Lu, Hisashi Kawai |
ICASSP | 2 |
| 2017 | Semi-supervised ensemble DNN acoustic model trainingabstractIt is very important to exploit abundant unlabeled speech for improving the acoustic model training in automatic speech recognition (ASR). Semi-supervised training methods incorporate unlabeled data in addition to labeled data to enhance the model training, but it encounters the error-prone label problem. The ensemble training scheme trains a set of models and combines them to make the model more general and robust, but it has not been applied to the unlabeled data. In this work, we propose an effective semi-supervised training of deep neural network (DNN) acoustic models by incorporating the diversity among the ensemble of models. The resultant model improved the performance in the lecture transcription task. Moreover, the proposed method has also shown a potential for DNN adaptation. Sheng Li 0010, Xugang Lu, Shinsuke Sakai, Masato Mimura, Tatsuya Kawahara |
ICASSP | 2 |
| 2017 | Conditional Generative Adversarial Nets Classifier for Spoken Language Identification
Xugang Lu, Sheng Li 0010, Hisashi Kawai |
INTERSPEECH | 2 |
| 2017 | Regularization of neural network model with distance metric learning for i-vector based spoken language identification
Xugang Lu, Yu Tsao 0001, Hisashi Kawai |
Comput. Speech Lang. | 1 |
| 2017 | Maximum-a-Posteriori-Based Decoding for End-to-End Acoustic ModelsabstractThis paper presents a novel decoding framework for acoustic models (AMs) based on end-to-end neural networks (e.g., connectionist temporal classification). The end-to-end training of AMs has recently demonstrated high accuracy and efficiency in automatic speech recognition (ASR). When using the trained AM in decoding, although a language model (LM) is implicitly involved in such an end-to-end AM, it is still essential to integrate an external LM trained with a large text corpus to achieve the best results. While there is no theoretical justification, most of the studies suggest using a naive interpolation of the end-to-end AM score and the external LM score, empirically. In this paper, we propose a more theoretically sound decoding framework derived from a maximization of the posterior probability of a word sequence given an observation. As a consequence of the theory, the subword LM is newly introduced to seamlessly integrate the external LM score with the end-to-end AM score. Our proposed method can be achieved by a small modification of the conventional weighted finite-state transducer-based implementation, without having to heavily increase the graph size. We tested the proposed decoding framework on ASR experiments with the Corpus of the Wall Street Journal and the Corpus of Spontaneous Japanese. The results showed that the proposed framework achieved significant and consistent improvements over the conventional interpolation-based decoding framework. Naoyuki Kanda, Xugang Lu, Hisashi Kawai |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Bottleneck linear transformation network adaptation for speaker adaptive training-based hybrid DNN-HMM speech recognizerabstractRecently, a Hybrid DNN-HMM recognizer trained with the Speaker Adaptive Training (SAT) concept was successfully modified to a more effective speaker-adaptation-oriented recognizer whose DNN front-end adopted a Linear Transformation Network (LTN) Speaker Dependent (SD) module. However, the size of SD modules is still large, which incurs high storage costs and the risk of over-training. To alleviate this problem, we analyze the characteristics of an LTN module by focusing on the relation between its size and its feature-representation capability. Moreover, we propose a new SAT-based scheme for reducing the LTN size using SVD-based matrix compression. Evaluation experiments on the TED Talks corpus prove that our LTN size-reduction scheme not only maintains the adaptation performance of the original LTN-embedded, SAT-based DNN-HMM recognizer but also further increases it especially in cases where the speech data available for adaptation training are severely limited. Tsubasa Ochiai, Shigeki Matsuda, Hideyuki Watanabe, Xugang Lu, Hisashi Kawai, Shigeru Katagiri |
ICASSP | 4 |
| 2016 | Local fisher discriminant analysis for spoken language identificationabstractI-vector is a state-of-the-art technique widely used in spoken language identification systems. Since i-vectors include total variability factors, discriminant analysis methods have been introduced to find the most discriminative features while removing the undesired variables for language identification, for example, linear discriminant analysis (LDA) and nonparametric discriminant analysis (NDA). However, these methods either do not consider or use weak local structures of the data. In this study, we introduce a local Fisher discriminant analysis (LFDA) as a post-processing discriminant analysis method to extract the discriminative features from i-vectors. LFDA is a full-rank method which takes the local structure of the data into account for non-Gaussian distribution data, i.e., multimodal. Compared with LDA and NDA, LFDA is a pair-wise local method which enhances the centralization of the distribution of samples in the same class to obtain larger amounts of discriminative features. Experimental results indicate that LFDA is more effective than LDA and NDA for the i-vector-based language identification task. Xugang Lu, Lemao Liu, Hisashi Kawai |
ICASSP | 2 |
| 2016 | SNR-Aware Convolutional Neural Network Modeling for Speech Enhancement
Szu-Wei Fu, Yu Tsao 0001, Xugang Lu |
INTERSPEECH | 3 |
| 2016 | Investigation of Semi-Supervised Acoustic Model Training Based on the Committee of Heterogeneous Neural Networks
Naoyuki Kanda, Shoji Harada, Xugang Lu, Hisashi Kawai |
INTERSPEECH | 3 |
| 2016 | Maximum a posteriori Based Decoding for CTC Acoustic Models
Naoyuki Kanda, Xugang Lu, Hisashi Kawai |
INTERSPEECH | 2 |
| 2016 | Pair-Wise Distance Metric Learning of Neural Network Model for Spoken Language Identification
Xugang Lu, Yu Tsao 0001, Hisashi Kawai |
INTERSPEECH | 1 |
| 2016 | F0 Contour Analysis Based on Empirical Mode Decomposition for DNN Acoustic Modeling in Mandarin Speech Recognition
Xiaoyun Wang 0002, Xugang Lu, Hisashi Kawai, Seiichi Yamamoto |
INTERSPEECH | 2 |
| 2016 | Combination of multiple acoustic models with unsupervised adaptation for lecture speech transcription
Xugang Lu, Xinhui Hu, Naoyuki Kanda, Masahiro Saiko, Chiori Hori, Hisashi Kawai |
Speech Commun. | 2 |
| 2016 | Wavelet Speech Enhancement Based on Nonnegative Matrix FactorizationabstractFor the state-of-the-art speech enhancement (SE) techniques, a spectrogram is usually preferred than the respective time-domain raw data, since it reveals more compact presentation together with conspicuous temporal information over a long time span. However, two problems can cause distortions in the conventional nonnegative matrix factorization (NMF)-based SE algorithms. One is related to the overlap-and-add operation used in the short-time Fourier transform (STFT)-based signal reconstruction, and the other is concerned with directly using the phase of the noisy speech as that of the enhanced speech in signal reconstruction. These two problems can cause information loss or discontinuity when comparing the clean signal with the reconstructed signal. To solve these two problems, we propose a novel SE method that adopts discrete wavelet packet transform (DWPT) and NMF. In brief, the DWPT is first applied to split a time-domain speech signal into a series of subband signals. Then, we exploit NMF to highlight the speech component for each subband. These enhanced subband signals are joined together via the inverse DWPT to reconstruct a noise-reduced signal in time domain. We evaluate the proposed DWPT-NMF-based SE method on the Mandarin hearing in noise test (MHINT) task. Experimental results show that this new method effectively enhances speech quality and intelligibility and outperforms the conventional STFT-NMF-based SE system. Syu-Siang Wang, Alan Chern, Yu Tsao 0001, Jeih-Weih Hung, Xugang Lu, Ying-Hui Lai, Borching Su |
IEEE Signal Process. Lett. | 5 |
| 2015 | Training data pseudo-shuffling and direct decoding framework for recurrent neural network based acoustic modelingabstractWe propose two techniques to enhance the performance of recurrent neural network (RNN)-based acoustic models. The first technique addresses training efficiency. Because RNNs require sequential input, it is difficult to randomly shuffle training samples to accelerate stochastic gradient descent based training. We propose a "pseudo-shuffling" procedure that instead augments training sample unexpectedness by skipping successive samples. The second proposed technique is a novel "direct decoding" framework in which the posterior probability of the RNN is inputted into a decoder without conversion into a hidden Markov model emission probability. In our large vocabulary speech recognition experiments with English lecture recordings, the first technique significantly improved RNN training efficiency, showing a 14.3% relative word error rate (WER) improvement. The second technique further achieved an additional 3.1% relative WER improvement. Our sigmoid-type RNN achieved a 10.7% better WER than same-sized deep neural networks without using long short-term memory cells. Naoyuki Kanda, Mitsuyoshi Tachimori, Xugang Lu, Hisashi Kawai |
ASRU | 3 |
| 2015 | Speaker adaptive training for deep neural networks embedding linear transformation networksabstractRecently, a novel speaker adaptation method was proposed that applied the Speaker Adaptive Training (SAT) concept to a speech recognizer consisting of a Deep Neural Network (DNN) and a Hidden Markov Model (HMM), and its utility was demonstrated. This method implements the SAT scheme by allocating one Speaker Dependent (SD) module for each training speaker to one of the intermediate layers of the front-end DNN. It then jointly optimizes the SD modules and the other part of network, which is shared by all the speakers. In this paper, we propose an improved version of the above SAT-based adaptation scheme for a DNN-HMM recognizer. Our new training adopts a Linear Transformation Network (LTN) for the SD module, and such LTN employment leads to more appropriate regularization in both the SAT and adaptation stages by replacing an empirically selected anchorage of a network for regularization in the preceding SAT-DNN-HMM with a SAT-optimized anchorage. We elaborate the effectiveness of our proposed method over TED Talks corpus data. Our experimental results show that a speaker-adapted recognizer using our method achieves a significant word error rate reduction of 9.2 points from a baseline SI-DNN recognizer and also steadily outperforms speaker-adapted recognizers, each of which originates from the preceding SAT-based DNN-HMM. Tsubasa Ochiai, Shigeki Matsuda, Hideyuki Watanabe, Xugang Lu, Chiori Hori, Shigeru Katagiri |
ICASSP | 4 |
| 2015 | Ensemble speaker modeling using speaker adaptive training deep neural network for speaker adaptationabstractIn this paper, we introduce an ensemble speaker modeling using a speaker adaptive training (SAT) deep neural network (SAT-DNN). We first train a speaker-independent DNN (SIDNN) acoustic model as a universal speaker model (USM). Based on the USM, a SAT-DNN is used to obtain a set of speaker-dependent models by assuming that all other layers except one speaker-dependent (SD) layer are shared among speakers. The speaker ensemble matrix is created by concatenating all of the SD neural weight matrices. With matrix factorization technique, an ensemble speaker subspace is extracted. When testing, an initial model for each target speaker is selected in this ensemble speaker subspace. Then, adaptation is carried out to obtain the final acoustic model for testing. In order to reduce the number of adaptation parameters, low-rank speaker subspace is further explored. We test our algorithm on lecture transcription task. Experimental results showed that our proposed method is effective for unsupervised speaker adaptation. Index Terms: speaker adaptation, deep neural networks, ensemble modeling, lecture transcription Sheng Li 0010, Xugang Lu, Yuya Akita, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2015 | Sparse representation with temporal max-smoothing for acoustic event detection
Xugang Lu, Yu Tsao 0001, Chiori Hori, Hisashi Kawai |
INTERSPEECH | 1 |
| 2015 | Ensemble environment modeling using affine transform group
Yu Tsao 0001, Payton Lin, Ting-Yao Hu, Xugang Lu |
Speech Commun. | 4 |
| 2014 | Speech enhancement using segmental nonnegative matrix factorizationabstractThe conventional NMF-based speech enhancement algorithm analyzes the magnitude spectrograms of both clean speech and noise in the training data via NMF and estimates a set of spectral basis vectors. These basis vectors are used to span a space to approximate the magnitude spectrogram of the noise-corrupted testing utterances. Finally, the components associated with the clean-speech spectral basis vectors are used to construct the updated magnitude spectrogram, producing an enhanced speech utterance. Considering that the rich spectral-temporal structure may be explored in local frequency and time-varying spectral patches, this study proposes a segmental NMF (SNMF) speech enhancement scheme to improve the conventional frame-wise NMF-based method. Two algorithms are derived to decompose the original nonnegative matrix associated with the magnitude spectrogram; the first algorithm is used in the spectral domain and the second algorithm is used in the temporal domain. When using the decomposition processes, noisy speech signals can be modeled more precisely, and spectrograms regarding the speech part can be constituted more favorably compared with using the conventional NMF-based method. Objective evaluations using perceptual evaluation of speech quality (PESQ) indicate that the proposed SNMF strategy increases the sound quality in noise conditions and outperforms the well-known MMSE log-spectral amplitude (LSA) estimation. Hao-Teng Fan, Jeih-Weih Hung, Xugang Lu, Syu-Siang Wang, Yu Tsao 0001 |
ICASSP | 3 |
| 2014 | Sparse representation based on a bag of spectral exemplars for acoustic event detectionabstractAcoustic event detection is an important step for audio content analysis and retrieval. Traditional detection techniques model the acoustic events on frame-based spectral features. Considering the temporal-frequency structures of acoustic events may be distributed in time-scales beyond frames, we propose to represent those structures as a bag of spectral patch exemplars. In order to learn the representative exemplars, k-means clustering based vector quantization (VQ) was applied on the whitened spectral patches which makes the learned exemplars focus on high-order statistical structure. With the learned spectral exemplars, a sparse feature representation is extracted based on the similarity measurement to the learned exemplars. A support vector machine (SVM) classifier was built on the sparse representation for acoustic event detection. Our experimental results showed that the sparse representation based on the patch based exemplars significantly improved the performance compared with traditional frame based representations. Xugang Lu, Yu Tsao 0001, Shigeki Matsuda, Chiori Hori |
ICASSP | 1 |
| 2014 | Speaker Adaptive Training using Deep Neural NetworksabstractAmong many speaker adaptation embodiments, Speaker Adaptive Training (SAT) has been successfully applied to a standard Hidden-Markov-Model (HMM) speech recognizer, whose state is associated with Gaussian Mixture Models (GMMs). On the other hand, recent studies on Speaker-Independent (SI) recognizer development have reported that a new type of HMM speech recognizer, which replaces GMMs with Deep Neural Networks (DNNs), outperforms GMM-HMM recognizers. Along these two lines, it is natural to conceive of further improvement to a preset DNN-HMM recognizer by employing SAT. In this paper, we propose a novel training scheme that applies SAT to a SI DNN-HMM recognizer. We then implement the SAT scheme by allocating a Speaker-Dependent (SD) module to one of the intermediate layers of a seven-layer DNN, and elaborate its utility over TED Talks corpus data. Experiment results show that our speaker-adapted SAT-based DNN-HMM recognizer reduces the word error rate by 8.4% more than that of a baseline SI DNN-HMM recognizer, and (regardless of the SD module allocation) outperforms the conventional speaker adaptation scheme. The results also show that the inner layers of DNN are more suitable for the SD module than the outer layers. Tsubasa Ochiai, Shigeki Matsuda, Xugang Lu, Chiori Hori, Shigeru Katagiri |
ICASSP | 3 |
| 2014 | Ensemble modeling of denoising autoencoder for speech spectrum restoration
Xugang Lu, Yu Tsao 0001, Shigeki Matsuda, Chiori Hori |
INTERSPEECH | 1 |
| 2014 | Incorporating local information of the acoustic environments to MAP-based feature compensation and acoustic model adaptationabstractThe maximum a posteriori (MAP) criterion is popularly used for feature compensation (FC) and acoustic model adaptation (MA) to reduce the mismatch between training and testing data sets. MAP-based FC and MA require prior densities of mapping function parameters, and designing suitable prior densities plays an important role in obtaining satisfactory performance. In this paper, we propose to use an environment structuring framework to provide suitable prior densities for facilitating MAP-based FC and MA for robust speech recognition. The framework is constructed in a two-stage hierarchical tree structure using environment clustering and partitioning processes. The constructed framework is highly capable of characterizing local information about complex speaker and speaking acoustic conditions. The local information is utilized to specify hyper-parameters in prior densities, which are then used in MAP-based FC and MA to handle the mismatch issue. We evaluated the proposed framework on Aurora-2, a connected digit recognition task, and Aurora-4, a large vocabulary continuous speech recognition (LVCSR) task. On both tasks, experimental results showed that with the prepared environment structuring framework, we could obtain suitable prior densities for enhancing the performance of MAP-based FC and MA. Yu Tsao 0001, Xugang Lu, Paul R. Dixon, Ting-Yao Hu, Shigeki Matsuda, Chiori Hori |
Comput. Speech Lang. | 2 |
| 2013 | Automatic localization of a language-independent sub-network on deep neural networks trained by multi-lingual speechabstractDeep neural networks (DNNs) have been successfully applied to automatic speech recognition (ASR). However, no study has investigated the possibility of building a language-independent sub-network DNN as the basis for further training of any new language using a simple plug-in of the sub-network. In this paper, we propose a novel technique to split a DNN into language-independent and -dependent sub-networks using multi-lingual speech training data. Our basic assumption is that, in a DNN for speech processing, language-independent feature processing is done in stages that are near to the input layer, while language-dependent processing is performed in stages that are near to the output layer. Based on this assumption, we propose a technique to simultaneously optimize multiple sub-networks in a DNN trained with multi-lingual speech data. The language-dependent and -independent processing boundaries in individual sub-networks are segmented automatically. We test our technique in phoneme classification experiments. The results demonstrate that a language-independent sub-network DNN extracted by our technique can be used as a universal network for speech processing of additional new languages. Shigeki Matsuda, Xugang Lu, Hideki Kashioka |
ICASSP | 2 |
| 2013 | Speech spectrum restoration based on conditional restricted boltzmann machine
Xugang Lu, Shigeki Matsuda, Chiori Hori |
INTERSPEECH | 1 |
| 2013 | Speech enhancement based on deep denoising autoencoderabstractWe previously have applied deep autoencoder (DAE) for noise reduction and speech enhancement. However, the DAE was trained using only clean speech. In this study, we further introduce an explicit denoising process in learning the DAE. In training the DAE, we still adopt greedy layer-wised pretraining plus fine tuning strategy. In pretraining, each layer is trained as a one hidden layer neural autoencoder (AE) using noisy-clean speech pairs as input and output (or transformed noisy-clean speech pairs by preceding AEs). Fine tuning was done by stacking all AEs with pretrained parameters for initialization. The trained DAE is used as a filter for speech estimation when noisy speech is given. Speech enhancement experiments were done to examine the performance of the trained denoising DAE. Noise reduction, speech distortion, and perceptual evaluation of speech quality (PESQ) criteria are used in the performance evaluations. Experimental results show that adding depth of the DAE consistently increase the performance when a large training data set is given. In addition, compared with a minimum mean square error based speech enhancement algorithm, our proposed denoising DAE provided superior performance on the three objective evaluations. Xugang Lu, Yu Tsao 0001, Shigeki Matsuda, Chiori Hori |
INTERSPEECH | 1 |
| 2012 | Factored Language Model based on Recurrent Neural Network
Youzheng Wu, Xugang Lu, Hitoshi Yamamoto, Shigeki Matsuda, Chiori Hori, Hideki Kashioka |
COLING | 2 |
| 2012 | Noise estimation using a constrained sequential HMM IN log-spectral domainabstractHow to utilize the time correlation of speech/nonspeech presence is a crucial problem faced by noise estimators. The popular technique of exploiting such correlation is to smooth noisy spectra by using a temporal recursive filter with a time-varying smoothing factor. But this technique cannot warrant the statistical optimality. In theory, hidden Markov model (HMM) is more desirable than this technique. It can give an elaborate description of speech/nonspeech transition. Moreover, some theoretical frameworks, such as maximum likelihood (ML), are available for optimal estimation. This paper presents a constrained sequential HMM to model the time correlation of speech/nonspeech presence of an individual log-power sequence. Its parameter set is on-line adapted to varying signals based on a ML framework. We compared its performance with that of well-established algorithms by speech enhancement experiments. The results confirmed its promising performance. Dongwen Ying, Xugang Lu, Yonghong Yan 0002, Jianwu Dang 0001, Frank K. Soong |
ICASSP | 2 |
| 2012 | Speech restoration based on deep learning autoencoder with layer-wised pretraining
Xugang Lu, Shigeki Matsuda, Chiori Hori, Hideki Kashioka |
INTERSPEECH | 1 |
| 2011 | Adaptive Regularization Framework for Robust Voice Activity Detection
Xugang Lu, Masashi Unoki, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura 0001 |
INTERSPEECH | 1 |
| 2011 | Voice Activity Detection in MTF-Based Power Envelope Restoration
Masashi Unoki, Xugang Lu, Rico Petrick, Shota Morita, Masato Akagi, Rüdiger Hoffmann |
INTERSPEECH | 2 |
| 2011 | Sub-band temporal modulation envelopes and their normalization for automatic speech recognition in reverberant environments
Xugang Lu, Masashi Unoki, Satoshi Nakamura 0001 |
Comput. Speech Lang. | 1 |
| 2011 | Temporal modulation normalization for robust speech feature extraction and recognition
Xugang Lu, Shigeki Matsuda, Masashi Unoki, Satoshi Nakamura 0001 |
Multim. Tools Appl. | 1 |
| 2010 | Voice activity detection in a reguarized reproducing kernel hilbert space
Xugang Lu, Masashi Unoki, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura 0001 |
INTERSPEECH | 1 |
| 2010 | Temporal contrast normalization and edge-preserved smoothing of temporal modulation structures of speech for robust speech recognition
Xugang Lu, Shigeki Matsuda, Masashi Unoki, Satoshi Nakamura 0001 |
Speech Commun. | 1 |
| 2010 | Vowel Production Manifold: Intrinsic Factor Analysis of Vowel ArticulationabstractThe manner in which the organization of vowels in a compact space reflects the relationship between their production and perception remains to be clarified in the field of speech science, although it is believed that vowels exist in a compact space with a regular structure rather than as an unstructured blob. In the articulatory domain, some traditional representations such as those based on the results of tongue position analysis and linear factor analysis are used. However, the former partially encodes information on vowel articulation, and the latter only reflects the linear degrees of freedom of vowel articulation. Since nonlinear degrees of freedom exist during the production of vowels, the traditional linear factor analysis is not suitable. In this paper, we proposed the use of Laplacian eigenmaps for analyzing the intrinsic factors affecting vowel articulation and obtain a compact manifold representation of vowels. On the manifold, vowels have distinct cluster positions depending on the similarities in their production and perception. On the basis of vowel articulation, we state that the first dimension of the manifold structure is related to the tongue height position and the second dimension to mouth opening. The third dimension is related to the articulation location of vowels along the vocal tract which is curved along the manifold. A similar topological manifold structure is explored in the vowel acoustic space. Quantitatively, on the basis of the conditional entropy criterion and vowel identification experiments, we confirmed that the analyzed compact manifold structure encodes more information on vowels than do traditional representations. Xugang Lu, Jianwu Dang 0001 |
IEEE Trans. Speech Audio Process. | 1 |
| 2009 | Temporal contrast normalization and edge-preserved smoothing on temporal modulation structure for robust speech recognitionabstractIn this paper, we propose a two-step processing algorithm which adaptively normalizes the temporal modulation of speech to extract robust speech feature for automatic speech recognition systems. The first step processing is to normalize the temporal modulation contrast (TMC) of the cepstral time series for both clean and noisy speech. The second step processing is to smooth the normalized temporal modulation structure to reduce the artifacts due to noise while preserving the speech modulation events (edges). We tested our algorithm on speech recognition experiments in additive noise condition (AURORA-2J data corpus), reverberant noise condition (convolution of clean speech utterances from AURORA-2J with a smart room impulse response), and noisy condition with both reverberant and additive noise (air conditioner noise in a smart room). For comparison, the ETSI advanced front-end (AFE) algorithm was used. Our results showed that the algorithm provided: (1) for additive noise condition, 57.26% relative word error reduction (RWER) rate for clean conditional training (59.37% for AFE), and 33.52% RWER rate for multi-conditional training (35.77% for AFE), (2) for reverberant condition, 51.28% RWER rate (10.17% for AFE) and (3) for noisy condition with both reverberant and additive noise, 71.74% RWER rate (48.86% for AFE). Xugang Lu, Shigeki Matsuda, Masashi Unoki, Tohru Shimizu, Satoshi Nakamura 0001 |
ICASSP | 1 |
| 2009 | Subband temporal modulation spectrum normalization for automatic speech recognition in reverberant environments
Xugang Lu, Masashi Unoki, Satoshi Nakamura 0001 |
INTERSPEECH | 1 |
| 2008 | A model based investigation of activation patterns of the tongue muscles for vowel production
Qiang Fang 0003, Satoru Fujita, Xugang Lu, Jianwu Dang 0001 |
INTERSPEECH | 3 |
| 2008 | Robust front end processing for speech recognition in reverberant environments: utilization of speech characteristicsabstractThis paper proposes two methods for robust automatic speech recognition (ASR) in reverberant environments. Unlike other methods which mostly apply inverse filtering by blindly estimated room impulse responses to achieve dereverberation, theproposed methods are based on the utilization of the characteristics of speech. The first method - Harmonicity based Feature Analysis – takes advantage of the harmonic componentsof speech, which are assumed to be undistorted. The second method - Temporal Power Envelope Feature Analysis – utilizes the temporal modulation structure of speech, representing the phoneme level temporal events which contain most intelligibility information. Both methods increase the recognition performance remarkably in a different way. Combining both of them connects their individual advantages. In order to examine theperformance of utilizing harmonicity and modulation temporal structure for reverberant ASR, the methods are tested in clean and reverberant training. As results show, even in strong reverberantconditions both methods obtain practical applicableperformance for reverberant training. In addition, besides testing their performance in dependency on the reverberation time, their performance considering the speaker-to-microphone distanceis tested, which is another new contributions in this paper. Rico Petrick, Xugang Lu, Masashi Unoki, Masato Akagi, Rüdiger Hoffmann |
INTERSPEECH | 2 |
| 2008 | An investigation of dependencies between frequency components and speaker characteristics for text-independent speaker identification
Xugang Lu, Jianwu Dang 0001 |
Speech Commun. | 1 |
| 2007 | Physiological Feature Extraction for Text Independent Speaker Identification using Non-Uniform Subband ProcessingabstractThe features used for speech recognition should emphasize linguistic information while suppressing speaker differences. For speaker recognition, features should have more speaker individual information while attenuating the linguistic information. In most studies, however, the identical acoustic features are used for the different missions of speaker and speech recognitions. In this paper, we propose a new physiological feature extraction method which emphasizes individual information for speaker identification. For the purpose, physiological features of speakers were analyzed from the point of view of speech production. It is found that the speaker individual information is encoded in different frequency regions of speech sound. The speaker discriminative information was quantified using Fisher's F-ratio in each frequency region. Based on the F-ratio, we proposed a non-uniform sub-band processing strategy to extract new feature which can emphasize or refine the physiological aspects involved in speech production. We combined the new feature with GMM for speaker identification task and applied on NTT-VR speaker recognition database. Compared with MFCC feature, by using the proposed feature, the identification error rate was reduced 20.1%. Xugang Lu, Jianwu Dang 0001 |
ICASSP (4) | 1 |
| 2007 | Dimension reduction for speaker identification based on mutual informationabstractAbstract Dimension reduction is a necessary step for speech feature extraction in a speaker identification system. Discrete Cosine Transform (DCT) or Principal Component Analysis (PCA) is widely used for dimension reduction. By choosing basis vectors from basis vector pool of DCT or PCA which contribute more to data distribution variance or reconstruction accuracy of speech data set, we can transform the data set by projecting them on to the selected basis vectors. However, keeping the maximum distribution variance or high reconstruction accuracy does not guarantee the optimal keeping of high speaker discriminative information. In this paper, we proposed a basis vector selection method based on mutual information concept which guarantees the keeping of high speaker discriminative information. The mutual information is used to measure the dependency between the features extracted using basis vectors and speaker class labels. The high mutual information related basis vectors are chosen for feature extraction. Considering one speaker feature may be encoded in more than one basis vectors, we proposed to use joint mutual information concept which takes the dependency between feature variables into consideration. Based on the selected basis vectors from DCT or PCA basis vector pool, we extracted features for speaker identification experiments. Experimental results showed that the speaker identification error rate using proposed feature was reduced 11% and 8% on average for DCT and PCA based features respectively. Xugang Lu, Jianwu Dang 0001 |
INTERSPEECH | 1 |
| 2006 | A robust feature extraction based on the MTF concept for speech recognition in reverberant environment
Xugang Lu, Masashi Unoki, Masato Akagi |
INTERSPEECH | 1 |
| 2006 | A simulation based parameter optimization for a coarticulation modelabstractA coarticulation model, namely ‘carrier model’, has been proposed previously by Dang et al. to improve the performance of a physiological articulatory model based speech synthesizer. The carrier model offers a good framework to account for coarticulation in the planning stage, while its parameters need to be refined for improving the performance of the model. This study is to refine the parameters of the carrier model and estimate typical phonetic targets by minimizing the differences between model simulations and observations. A simulation based optimization framework is proposed for this purpose. The framework consists of two layers: obtaining planned targets in a low layer; estimating phonetic targets and optimizing the parameters in a high layer. A direct search method was applied to the low layer due to the non-analytic nature of the articulation model, while the high layer adopts bilevel optimization strategy to decompose the complicated problem into a set of subproblems. A general evaluation was conducted by combining the refined carrier model and the learned phonetic targets together using the physiological articulatory model and the average error between observations and simulations was 0.15 cm over 103 VCV combinations on the jaw, tongue tip and tongue dorsum. Index Terms: speech production, coarticulation, optimization Jianguo Wei, Xugang Lu, Jianwu Dang 0001 |
INTERSPEECH | 2 |
| 2005 | A noise reduction system in arbitrary noise environments and its applications to speech enhancement and speech recognitionabstractThe paper proposes a novel noise reduction system in arbitrary noise environments, consisting of localized and nonlocalized noises, where few existing systems work well. In the proposed system, localized noises are estimated and reduced by the hybrid noise estimation technique we previously proposed and spectral subtraction. Non-localized noises are reduced by a post-filter whose performance is further improved by a novel estimator for the a priori speech absence probability calculated under the assumption of a diffuse noise field. Experimental results show that the proposed system results in significant improvements in terms of speech quality measures and speech recognition performance in various noise conditions. Xugang Lu, Masato Akagi |
ICASSP (3) | 2 |
| 2000 | Dominant subspace analysis for auditory spectrumabstractIn hearing perception theory, spectral structure is a most important feature for speech perception, this spectral structure is not easy to be masked in noisy condition. So if this structure is extracted and enhanced, the representation will be much more robust. In this paper, we propose a new statistical dominant subspace analysis method for auditory spectrum based on SVD(Singular Values Decomposition) and signal subspace analysis method. The auditory spectrum can be decomposed into two subspaces, one is a dominant subspace, which is expanded by useful speech auditory spectrum , another subspace is sub-dominant subspace, which there is only noise information. So we analysis the auditory spectrum in the dominant subspace, the SNR will be increased. Thus this representation is much more robust. 1. COMPUTATIONAL AUDITORY MODEL AND AUDITORY SPECTRUM Speech stimulation can be represented by auditory neural system in many stages. First, it can be decomposed into many frequency bands by basilar membrane, then after processed by inner hair cell and neural fibers, it's intensity is represented by neural firing rate. This neural impulse can be transformed to auditory central system, where it can be perceived by auditory cortex[1]. In this paper, all the processing parts are integrated using digital signal processing method, when speech signal is processed by this model, auditory feature can be gotten. The basic processing frame is as in Fig.1: the system is made up of six parts, that is, the high pass filtering of outer ear and middle ear, the band pass filtering of basilar membrane, nonlinear compression and half wave rectifying of inner hair cell, low pass filtering of neural fiber, energy detection of central system, Figure1 Auditory model for speech signal processing A mathematical model is designed to simulate this auditory function, as in Fig.2, a low pass filter is used to simulate the long temporal integration mechanism. The function of outer/middle ear can be simulated by high pass filter; band pass filters for basilar membrane; halfwave rectify for inner hair cell; low pass filter for neural fiber; energy detector and log compression for neural central, at last a DCT is used to get the feature vector. Figure2 The mathematical model for Figure 1 In these modules, short term adaptation and rapid adaptation of inner hair cell and neural fiber are not considered. Also, functions of temporal integration of neural central system and low pass filtering of neural fiber are integrated as a low pass filter. Energy detector is used for the intensity detection for each frequency channel. After processed by this model, a auditory feature is gotten. The feature can be used for training and testing. In this paper, we only focus on the auditory spectrum analysis, so the auditory spectrum can be got from the energy detector of Figure 2. Visual representations of FFT, LPC, and Auditory Spectrum are drawn for comparison in figure 3.( the spectrum of a Chinese sentence ). Figure3 Top is FFT spectrum, middle is LPC spectrum. Bottom is Auditory Spectrum(AS) . From Fig.3, it is clear that AS(Auditory Spectrum) is wide band spectrum, FFT is narrow band spectrum. AS spectrum can be regarded as a smoothed spectrum of FFT spectrum in hearing perception scale(in frequency domain). It is very clear that , speech representation by auditory system is a series of time-frequency patches, these patches are different from noise patch. We can regard the speech feature as a continuous time-frequency patch with regular structure. Noise patches is random and no-regular, so we hope subspace decomposition method can help use to separate noise and speech by this property. 2 SIGNAL SUBSPACE AND SVD Signal subspace analysis method is widely used in digital signal processing and pattern recognition[2]. It is supposed that the useful information is only related with some lower dimensional subspaces, but noise is uniformly distributed in the whole measurement space (the whole Euclidean space). The subspace analysis method can decompose the whole measurement space into some useful subspaces, such as signal subspace and noise subspace or dominant subspace and subordinate subspace, then when the original feature is projected into the dominant subspaces, the dominant structure will be retained only, that is to say the subordinate feature(which including noise structure) will be reduced . From transformation view, we hope to find a new transform basis, the new feature gotten by this transformation can possess certain property, such as, each dimension of the feature vector is un-correlated or independent, etc. In this paper, we propose the subspace analysis method for the processing purpose. Singular Value Decomposition (SVD) is a very useful method for matrix structure decomposition. we give some useful formulas here which can be used later. Suppose the data matrix is n m R X × ∈ , there exist orthogonal matrices: m m m R u u U × ∈ = ) ,..., ( 1 n n m R v v V × ∈ = ) ,..., ( 1 satisfying: Xugang Lu, Lipo Wang 0001 |
INTERSPEECH | 1 |
| 1999 | Integrating spatial and temporal mechanisms in auditory neural fiber's computational modelabstractIn traditional speech signal processing methods and current auditory based methods, features are extracted based on power spectrum, that is, spatial or temporal mechanism is used to simulate the frequency response of our cochlear function. The disadvantage of these methods are that noise and tone signals are processed equally, but, in fact, our auditory system percepts noise and periodic stimulation with different sensitivity: if the stimulation is noise, the audible threshold is high, and the gain for noise is low. On the contrary, if the stimulation is periodic time series, then the auditory system's audible threshold will be low and the gain will be high, that is the temporal processing aspect. In this paper, spatial and temporal mechanisms are integrated in neural firing response, thus the representation not only represents the average firing rate of neural fibers, but also enhances the periodic components of the stimulation. Thus, this representation can have both merits of the two processing methods. Xugang Lu, Daowen Chen |
IJCNN | 1 |