EDBT 2026 Demo / reviewers in the wild / expert
Jun Qi 0002
dblp:133/4051-2
· DBLP profile ↗
24ranked-venue papers
16as first author
11since 2021 · last 2025
0000-0001-7533-2630ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 12 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 5 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Projection Valued-based Quantum Machine Learning Adapting to Differential Privacy Algorithm for Word-level LipreadingabstractDeep neural network (DNN)-based lipreading models have achieved excellent recognition accuracy but are currently facing challenges related to user privacy. To address this, we propose a novel hybrid quantum-classical neural network (HQCNN) for lipreading that balances superior performance with enhanced privacy protection. The HQCNN-based lipreading model features an innovative variational quantum circuit (VQC) back-end, which transforms the output of the DNN front-end into quantum representations and predicts the posterior probability of each word. Furthermore, we introduce projection-valued encoding (PVE) and projection-valued measurement (PVM), enabling the VQC to handle inputs and outputs of dimensions that scale exponentially with the number of qubits, thereby substantially increasing its expressive power. Additionally, we explore the privacy-preserving properties of the HQCNN-based lipreading model by integrating differentially private stochastic gradient descent (DP-SGD). Experiments conducted on the LRW dataset demonstrate the model’s exceptional recognition accuracy and privacy-preserving capabilities. Hang Chen 0001, Jun Du 0002, Chao-Han Huck Yang, Jun Qi 0002 |
ICASSP | 5 |
| 2025 | Quantum Machine Learning: An Interplay Between Quantum Computing and Machine LearningabstractQuantum machine learning (QML) is a rapidly growing field that combines quantum computing principles with traditional machine learning. It seeks to revolutionize machine learning by harnessing the unique capabilities of quantum mechanics and employs machine learning techniques to advance quantum computing research. This paper presents an overview of quantum computing for the machine learning paradigm, where variational quantum circuits (VQC) are used to develop QML architectures on noisy intermediate-scale quantum (NISQ) devices. We discuss machine learning for the quantum computing paradigm, showcasing our recent theoretical and empirical findings. In particular, we delve into future directions for studying QML, exploring the potential industrial impacts of QML research. Jun Qi 0002, Chao-Han Huck Yang, Samuel Yen-Chi Chen |
ISCAS | 1 |
| 2024 | Exploiting A Quantum Multiple Kernel Learning Approach For Low-Resource Spoken Command RecognitionabstractWe propose a theoretical analysis of quantum projection learning (QPL) that employs multiple kernels, highlighting its advantages through representation error analysis. Building upon previous studies that utilized a single quantum kernel-based method, we further investigate a quantum projection framework that incorporates multiple Gaussian kernels for low-resource spoken command recognition. Our empirical results align with our theoretical insights, suggesting that methods based on multiple kernels can further enhance the performance of QPL. By leveraging the quantum-to-classical projected output embeddings, we integrate this with a prototypical network for acoustic modeling. When evaluated using Arabic, Chuvash, Irish, and Lithuanian low-resource speech from CommonVoice, our proposed method surpasses the recurrent neural network and single kernel-based classifier baselines by an average of +5.28%. Xianyan Fu, Xiao-Lei Zhang 0001, Chao-Han Huck Yang, Jun Qi 0002 |
ICASSP | 4 |
| 2024 | Interpretable Spectrum Transformation Attacks to Speaker Recognition SystemsabstractThe success of adversarial attacks on speaker recognition is mainly in white-box scenarios. When applying the adversarial voices that are generated by attacking white-box surrogate models to black-box victim models, i.e. transfer-based black-box attacks, the transferability of the adversarial voices is not only far from satisfactory, but also lacks interpretable basis. To address these issues, in this paper, we propose a general framework, named spectral transformation attack based on modified discrete cosine transform (STA-MDCT), to improve the transferability of the adversarial voices to a black-box victim model. Specifically, we first apply MDCT to the input voice. Then, we slightly modify the energy of different frequency bands for capturing the salient regions of the adversarial noise in the time-frequency domain that are critical to a successful attack. Unlike existing approaches that operate voices in the time domain, the proposed framework operates voices in the time-frequency domain, which improves the interpretability, transferability, and imperceptibility of the attack. Moreover, it can be implemented with any gradient-based attackers. To utilize the advantage of model ensembling, we not only implement STA-MDCT with a single white-box surrogate model but also with an ensemble of surrogate models. Finally, we visualize the saliency maps of adversarial voices by the class activation maps (CAM), which offer an interpretable basis for transfer-based attacks in speaker recognition for the first time. Extensive comparison results with six representative attackers show that the CAM visualization clearly explains the effectiveness of STA-MDCT and the weaknesses of the comparison methods; the proposed method outperforms the comparison methods by a large margin. Our audio samples are available on the demo website. Jiadi Yao, Jun Qi 0002, Xiao-Lei Zhang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Optimizing Quantum Federated Learning Based on Federated Quantum Natural Gradient DescentabstractQuantum federated learning (QFL) is a quantum extension of the classical federated learning model across multiple local quantum devices. An efficient optimization algorithm is always expected to minimize the communication overhead among different quantum participants. In this work, we propose an efficient optimization algorithm, namely federated quantum natural gradient descent (FQNGD), and further, apply it to a QFL framework that is com-posed of a variational quantum circuit (VQC)-based quantum neural networks (QNN). Compared with stochastic gradient descent methods like Adam and Adagrad, the FQNGD algorithm admits much fewer training iterations for the QFL to get converged. Moreover, it can significantly reduce the total communication overhead among local quantum devices. Our experiments on a handwritten digit classification dataset justify the effectiveness of the FQNGD for the QFL framework in terms of a faster convergence rate on the training set and higher accuracy on the test set. Jun Qi 0002, Xiao-Lei Zhang 0001, Javier Tejedor |
ICASSP | 1 |
| 2023 | Exploiting Low-Rank Tensor-Train Deep Neural Networks Based on Riemannian Gradient Descent With Illustrations of Speech ProcessingabstractThis work focuses on designing low-complexity hybrid tensor networks by considering trade-offs between the model complexity and practical performance. Firstly, we exploit a low-rank tensor-train deep neural network (TT-DNN) to build an end-to-end deep learning pipeline, namely LR-TT-DNN. Secondly, a hybrid model combining LR-TT-DNN with a convolutional neural network (CNN), which is denoted as CNN+(LR-TT-DNN), is set up to boost the performance. Instead of randomly assigning large TT-ranks for TT-DNN, we leverage Riemannian gradient descent to determine a TT-DNN associated with small TT-ranks. Furthermore, CNN+(LR-TT-DNN) consists of convolutional layers at the bottom for feature extraction and several TT layers at the top to solve regression and classification problems. We separately assess the LR-TT-DNN and CNN+(LR-TT-DNN) models on speech enhancement and spoken command recognition tasks. Our empirical evidence demonstrates that the LR-TT-DNN and CNN+(LR-TT-DNN) models with fewer model parameters can outperform the TT-DNN and CNN+(TT-DNN) counterparts. Jun Qi 0002, Chao-Han Huck Yang, Javier Tejedor |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | Mitigating Clipping Distortion in Multicarrier Transmissions Using Tensor-Train Deep Neural NetworksabstractMulticarrier transmissions, such as orthogonal frequency/chirp division multiplexing (OF/CDM), offer high spectral efficiency and low complexity equalization in multipath fading channels at the cost of high peak-to-average power ratio (PAPR). High peak powers can occur randomly and may drive the power amplifier (PA) into saturation, resulting in non-linear distortion. In this work, we propose a novel tensor-train (TT) deep neural network (DNN) architecture combined with soft clipping to reduce the PAPR to a deterministic level and minimize in-band distortion, while also satisfying spectral mask constraints. The proposed solution requires modifications only at the transmitter and there is no loss of spectral efficiency. Employing the TT decomposition allows for significant reduction in parameters. Results show that the proposed solution allows for significant reduction in PAPR with minimal performance loss. Furthermore, an upper bound for the PAPR is also derived which allows for the prediction of the required input power back off (IBO) without extensive simulations and trial-and-error. Muhammad Shahmeer Omar, Jun Qi 0002, Xiaoli Ma |
IEEE Trans. Wirel. Commun. | 2 |
| 2022 | Exploiting Hybrid Models of Tensor-Train Networks For Spoken Command RecognitionabstractThis work aims to design a low complexity spoken command recognition (SCR) system by considering different trade-offs between the number of model parameters and classification accuracy. More specifically, we exploit a deep hybrid architecture of a tensor-train (TT) network to build an end-to-end SRC pipeline. Our command recognition system, namely CNN+(TT-DNN), is composed of convolutional layers at the bottom for spectral feature extraction and TT layers at the top for command classification. Compared with a traditional end-to-end CNN baseline for SCR, our proposed CNN+(TTDNN) model replaces fully connected (FC) layers with TT ones and it can substantially reduce the number of model parameters while maintaining the baseline performance of the CNN model. We initialize the CNN+(TT-DNN) model in a randomized manner or based on a well-trained CNN+DNN, and assess the CNN+(TT-DNN) models on the Google Speech Command Dataset. Our experimental results show that the proposed CNN+(TT-DNN) model attains a competitive accuracy of 96.31% with 4 times fewer model parameters than the CNN model. Furthermore, the CNN+(TT-DNN) model can obtain a 97.2% accuracy when the number of parameters is increased. Jun Qi 0002, Javier Tejedor |
ICASSP | 1 |
| 2022 | Classical-To-Quantum Transfer Learning for Spoken Command Recognition Based on Quantum Neural NetworksabstractThis work investigates an extension of transfer learning applied in machine learning algorithms to the emerging hybrid end-to-end quantum neural network (QNN) for spoken command recognition (SCR). Our QNN-based SCR system is composed of classical and quantum components: (1) the classical part mainly relies on a 1D convolutional neural network (CNN) to extract speech features; (2) the quantum part is built upon the variational quantum circuit with a few learnable parameters. Since it is inefficient to train the hybrid end-to-end QNN from scratch on a noisy intermediate-scale quantum (NISQ) device, we put forth a hybrid transfer learning algorithm that allows a pre-trained classical network to be transferred to the classical part of the hybrid QNN model. The pre-trained classical network is further modified and augmented through jointly fine-tuning with a variational quantum circuit (VQC). The hybrid transfer learning methodology is particularly attractive for the task of QNN-based SCR because low-dimensional classical features are expected to be encoded into quantum states. We assess the hybrid transfer learning algorithm applied to the hybrid classical-quantum QNN for SCR on the Google speech command dataset, and our classical simulation results suggest that the hybrid transfer learning can boost our baseline performance on the SCR task. Jun Qi 0002, Javier Tejedor |
ICASSP | 1 |
| 2022 | When BERT Meets Quantum Temporal Convolution Learning for Text Classification in Heterogeneous ComputingabstractThe rapid development of quantum computing has demonstrated many unique characteristics of quantum advantages, such as richer feature representation and more secured protection on model parameters. This work proposes a vertical federated learning architecture based on variational quantum circuits to demonstrate the competitive performance of a quantum-enhanced pre-trained BERT model for text classification. In particular, our proposed hybrid classical-quantum model consists of a novel random quantum temporal convolution (QTC) learning framework replacing some layers in the BERT-based decoder. Our experiments on intent classification show that our proposed BERT-QTC model attains competitive experimental results in the Snips and ATIS spoken language datasets. Particularly, the BERT-QTC boosts the performance of the existing quantum circuit-based language model in two text classification datasets by 1.57% and 1.52% relative improvements. Furthermore, BERT-QTC can be feasibly deployed on both existing commercial-accessible quantum computation hardware and CPU-based interface for ensuring data isolation. Chao-Han Huck Yang, Jun Qi 0002, Samuel Yen-Chi Chen, Yu Tsao 0001 |
ICASSP | 2 |
| 2021 | Decentralizing Feature Extraction with Quantum Convolutional Neural Network for Automatic Speech RecognitionabstractWe propose a novel decentralized feature extraction approach in federated learning to address privacy-preservation issues for speech recognition. It is built upon a quantum convolutional neural network (QCNN) composed of a quantum circuit encoder for feature extraction, and a recurrent neural network (RNN) based end-to-end acoustic model (AM). To enhance model parameter protection in a decentralized architecture, an input speech is first up-streamed to a quantum computing server to extract Mel-spectrogram, and the corresponding convolutional features are encoded using a quantum circuit algorithm with random parameters. The encoded features are then down-streamed to the local RNN model for the final recognition. The proposed decentralized framework takes advantage of the quantum learning progress to secure models and to avoid privacy leakage attacks. Testing on the Google Speech Commands Dataset, the proposed QCNN encoder attains a competitive accuracy of 95.12% in a decentralized model, which is better than the previous architectures using centralized RNN models with convolutional features. We also conduct an in-depth study of different quantum circuit encoder architectures to provide insights into designing QCNN-based feature extractors. Neural saliency analyses demonstrate a correlation between the proposed QCNN features, class activation maps, and input spectrograms. We provide an implementation for future studies. Chao-Han Huck Yang, Jun Qi 0002, Samuel Yen-Chi Chen, Sabato Marco Siniscalchi, Xiaoli Ma, Chin-Hui Lee 0001 |
ICASSP | 2 |
| 2020 | Tensor-To-Vector Regression for Multi-Channel Speech Enhancement Based on Tensor-Train NetworkabstractWe propose a tensor-to-vector regression approach to multi-channel speech enhancement in order to address the issue of input size explosion and hidden-layer size expansion. The key idea is to cast the conventional deep neural network (DNN) based vector-to-vector regression formulation under a tensor-train network (TTN) framework. TTN is a recently emerged solution for compact representation of deep models with fully connected hidden layers. Thus TTN maintains DNN's expressive power yet involves a much smaller amount of trainable parameters. Furthermore, TTN can handle a multi-dimensional tensor input by design, which exactly matches the desired setting in multi-channel speech enhancement. We first provide a theoretical extension from DNN to TTN based regression. Next, we show that TTN can attain speech enhancement quality comparable with that for DNN but with much fewer parameters, e.g., a reduction from 27 million to only 5 million parameters is observed in a single-channel scenario. TTN also improves PESQ over DNN from 2.86 to 2.96 by slightly increasing the number of trainable parameters. Finally, in 8-channel conditions, a PESQ of 3.12 is achieved using 20 million parameters for TTN, whereas a DNN with 68 million parameters can only attain a PESQ of 3.06. Jun Qi 0002, Hu Hu, Yannan Wang, Chao-Han Huck Yang, Sabato Marco Siniscalchi, Chin-Hui Lee 0001 |
ICASSP | 1 |
| 2020 | Submodular Rank Aggregation on Score-Based Permutations for Distributed Automatic Speech RecognitionabstractDistributed automatic speech recognition (ASR) requires to aggregate outputs of distributed deep neural network (DNN)-based models. This work studies the use of submodular functions to design a rank aggregation on score-based permutations, which can be used for distributed ASR systems in both supervised and unsupervised modes. Specifically, we compose an aggregation rank function based on the Lovasz Bregman divergence for setting up linear structured convex and nested structured concave functions. The algorithm is based on stochastic gradient descent (SGD) and can obtain well-trained aggregation models. Our experiments on the distributed ASR system show that the submodular rank aggregation can obtain higher speech recognition accuracy than traditional aggregation methods like Adaboost. Code is available online1. Jun Qi 0002, Chao-Han Huck Yang, Javier Tejedor |
ICASSP | 1 |
| 2020 | Characterizing Speech Adversarial Examples Using Self-Attention U-Net EnhancementabstractRecent studies have highlighted adversarial examples as ubiquitous threats to the deep neural network (DNN) based speech recognition systems. In this work, we present a U-Net based attention model, UNetAt, to enhance adversarial speech signals. Specifically, we evaluate the model performance by interpretable speech recognition metrics and discuss the model performance by the augmented adversarial training. Our experiments show that our proposed U-NetAtimproves the perceptual evaluation of speech quality (PESQ) from 1.13 to 2.78, speech transmission index (STI) from 0.65 to 0.75, shortterm objective intelligibility (STOI) from 0.83 to 0.96 on the task of speech enhancement with adversarial speech examples. We conduct experiments on the automatic speech recognition (ASR) task with adversarial audio attacks. We find that (i) temporal features learned by the attention network are capable of enhancing the robustness of DNN based ASR models; (ii) the generalization power of DNN based ASR model could be enhanced by applying adversarial training with an additive adversarial data augmentation. The ASR metric on word-error-rates (WERs) shows that there is an absolute 2.22 % decrease under gradient-based perturbation, and an absolute 2.03 % decrease, under evolutionary-optimized perturbation, which suggests that our enhancement models with adversarial training can further secure a resilient ASR system. Chao-Han Huck Yang, Jun Qi 0002, Xiaoli Ma, Chin-Hui Lee 0001 |
ICASSP | 2 |
| 2020 | Enhanced Adversarial Strategically-Timed Attacks Against Deep Reinforcement LearningabstractRecent deep neural networks based techniques, especially those equipped with the ability of self-adaptation in the system level such as deep reinforcement learning (DRL), are shown to possess many advantages of optimizing robot learning systems (e.g., autonomous navigation and continuous robot arm control.) However, the learning-based systems and the associated models may be threatened by the risks of intentionally adaptive (e.g., noisy sensor confusion) and adversarial perturbations from real-world scenarios. In this paper, we introduce timing-based adversarial strategies against a DRL-based navigation system by jamming in physical noise patterns on the selected time frames. To study the vulnerability of learning-based navigation systems, we propose two adversarial agent models: one refers to online learning; another one is based on evolutionary learning. Besides, three open-source robot learning and navigation control environments are employed to study the vulnerability under adversarial timing attacks. Our experimental results show that the adversarial timing attacks can lead to a significant performance drop, and also suggest the necessity of enhancing the robustness of robot learning systems. Chao-Han Huck Yang, Jun Qi 0002, I-Te Danny Hung, Chin-Hui Lee 0001, Xiaoli Ma |
ICASSP | 2 |
| 2020 | Exploring Deep Hybrid Tensor-to-Vector Network Architectures for Regression Based Speech EnhancementabstractThis paper investigates different trade-offs between the number of model parameters and enhanced speech qualities by employing several deep tensor-to-vector regression models for speech enhancement. We find that a hybrid architecture, namely CNN-TT, is capable of maintaining a good quality performance with a reduced model parameter size. CNN-TT is composed of several convolutional layers at the bottom for feature extraction to improve speech quality and a tensor-train (TT) output layer on the top to reduce model parameters. We first derive a new upper bound on the generalization power of the convolutional neural network (CNN) based vector-to-vector regression models. Then, we provide experimental evidence on the Edinburgh noisy speech corpus to demonstrate that, in single-channel speech enhancement, CNN outperforms DNN at the expense of a small increment of model sizes. Besides, CNN-TT slightly outperforms the CNN counterpart by utilizing only 32% of the CNN model parameters. Besides, further performance improvement can be attained if the number of CNN-TT parameters is increased to 44% of the CNN model size. Finally, our experiments of multi-channel speech enhancement on a simulated noisy WSJ0 corpus demonstrate that our proposed hybrid CNN-TT architecture achieves better results than both DNN and CNN models in terms of better-enhanced speech qualities and smaller parameter sizes. Jun Qi 0002, Hu Hu, Yannan Wang, Chao-Han Huck Yang, Sabato Marco Siniscalchi, Chin-Hui Lee 0001 |
INTERSPEECH | 1 |
| 2020 | On Mean Absolute Error for Deep Neural Network Based Vector-to-Vector RegressionabstractIn this paper, we exploit the properties of mean absolute error (MAE) as a loss function for the deep neural network (DNN) based vector-to-vector regression. The goal of this work is two-fold: (i) presenting performance bounds of MAE, and (ii) demonstrating new properties of MAE that make it more appropriate than mean squared error (MSE) as a loss function for DNN based vector-to-vector regression. First, we show that a generalized upper-bound for DNN-based vector-to-vector regression can be ensured by leveraging the known Lipschitz continuity property of MAE. Next, we derive a new generalized upper bound in the presence of additive noise. Finally, in contrast to conventional MSE commonly adopted to approximate Gaussian errors for regression, we show that MAE can be interpreted as an error modeled by Laplacian distribution. Speech enhancement experiments are conducted to corroborate our proposed theorems and validate the performance advantages of MAE over MSE for DNN based regression. Jun Qi 0002, Jun Du 0002, Sabato Marco Siniscalchi, Xiaoli Ma, Chin-Hui Lee 0001 |
IEEE Signal Process. Lett. | 1 |
| 2019 | A Theory on Deep Neural Network Based Vector-to-Vector Regression With an Illustration of Its Expressive Power in Speech EnhancementabstractThis paper focuses on a theoretical analysis of deep neural network (DNN) based functional approximation. Leveraging upon two classical theorems on universal approximation, an artificial neural network (ANN) with a single hidden layer of neurons is used. With modified ReLU and Sigmoid activation functions, we first generalize the related concepts to vector-to-vector regression. Then, we show that the width of the hidden layer of ANN is numerically related to the approximation of the regression function. Furthermore, we increase the number of hidden layers and show that the depth of the ANN-based regression function can enhance its expressive power. We illustrate this representation with recently-emerged DNN based speech enhancement. We first compare the expressive power by varying ANN structures and then test its related regression performance under different noisy conditions in various noise types and signal-to-noise-ratio levels. Experimental results verify our theoretical prediction that an ANN of a broader hidden layer and a deeper architecture can jointly ensure a closer approximation of the vector-to-vector regression functions in terms of the Euclidean distance between the log power spectra of noisy and expected clean speech. Moreover, a DNN with a broader width at the top hidden layer can improve the regression performance relative to those with a narrower width at the top hidden layers. Jun Qi 0002, Jun Du 0002, Sabato Marco Siniscalchi, Chin-Hui Lee 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2018 | Distributed Submodular Maximization for Large Vocabulary Continuous Speech RecognitionabstractHuge training datasets for automatic speech recognition (ASR) typically contain redundant information so that a subset of data is generally enough to obtain similar ASR performance to that obtained when the entire dataset is employed for training. Although the centralized submodular-based data selection methods have been successfully applied to obtain a representable subset involving the most significant information of the whole dataset, the submodular data selection conveys problems in adapting to an extremely massive dataset. This paper proposes to use distributed submodular maximization (DSM) for efficiently selecting a data subset that maintains the ASR performance, while reducing tremendously the computational overhead. There are two approaches for the distributed submodular maximization problem: one is based on an homogeneous submodular function, and the other relies on decomposable submodular functions in which heterogeneous submodular functions are applied. Our experiments show that the data subset output by the DSM algorithms can maintain the ASR performance, while significantly reducing the computational overhead.1 Jun Qi 0002, Xu Liu 0017, Shunsuke Kamijo, Javier Tejedor |
ICASSP | 1 |
| 2016 | Deep multi-view representation learning for multi-modal features of the schizophrenia and schizo-affective disorderabstractThis work is originated from the MLSP 2014 Classification Challenge which tries to automatically detect subjects with schizophrenia and schizo-affective disorder by analyzing multi-modal features derived from magnetic resonance imaging (MRI) data. We employ Deep Neural Network (DNN)-based multi-view representation learning for combining multimodal features. The DNN-based multi-view models include deep canonical correlation analysis (DCCA) and deep canonically correlated auto-encoders (DCCAE). In addition, support vector machine with Gaussian kernel is used to conduct classification with the compact bottleneck features learned by the deep multi-view models. Our experiments on the dataset provided by the MLSP Classification Challenge show that bottleneck features learned via deep multi-view models obtain better results than the trimming features used in the baseline system in terms of the receiver operating characteristic (ROC) area under the curve (AUC). Jun Qi 0002, Javier Tejedor |
ICASSP | 1 |
| 2016 | Robust submodular data partitioning for distributed speech recognitionabstractDistributed deep neural networks are commonly employed for building automatic speech recognition (ASR) systems. In this work, we employ the robust submodular partitioning approach, which aims to split the training data into small disjoint data subsets and use each of these subsets to train a particular deep neural network. Two efficient algorithms are used as robust submodular functions [1], namely `Greedi-Max' and `Minorization-Maximization' [2], which are guaranteed to provide tight approximations to the submodular data partition problem. Experiments on TIMIT database show that each of the distributed neural networks trained by the submodular data subset obtains better results than that trained on any subset of data partitioned in a random way., In addition, multi-class adaboost is effectively used to fuse the outputs of the deep neural networks and provides competitive ASR results compared with the traditional ASR system. Besides, the time incurred by acoustic modeling is significantly reduced, which delivers us further benefits. Jun Qi 0002, Javier Tejedor |
ICASSP | 1 |
| 2013 | Subspace models for bottleneck featuresabstractThe bottleneck (BN) feature, particularly based on deep structures, has gained significant success in automatic speech recognition (ASR). However, applying the BN feature to small/medium-scale tasks is nontrivial. An obvious reason is that the limited training data prevent from training a complicated deep network; another reason, which is more subtle, is that the BN feature tends to possess high inter-dimensional correlation, thus being inappropriate to be modeled by the conventional diagonal Gaussian mixture model (GMM). This difficulty can be mitigated by increasing the number of Gaussian components and/or employing full covariance matrices. These approaches, however, are not applicable for small/medium-scale tasks for which only a limited amount of training data is available. In this paper, we study the subspace Gaussian mixture model (SGMM) for BN features. The SGMM assumes full but shared covariance matrices, and hence can address the interdimensional correlation in a parsimonious way. This is particularly attractive for the BN feature, especially on small/mediumscale tasks, where the inter-dimensional correlation is high but the full covariance modeling is not affordable due to the limited training data. Our preliminary experiments on the Resource Management (RM) database demonstrate that the SGMM can deliver significant performance improvement for ASR systems based on BN features. Jun Qi 0002, Dong Wang 0013, Javier Tejedor |
INTERSPEECH | 1 |
| 2013 | Bottleneck features based on gammatone frequency cepstral coefficientsabstractRecent work demonstrates impressive success of the bottleneck (BN) feature in speech recognition, particularly with deep networks plus appropriate pre-training. A widely admitted advantage associated with the BN feature is that the network structure can learn multiple environmental conditions with abundant training data. For tasks with limited training data, however, this multi-condition training is unavailable, and so the networks tend to be over-fitted and sensitive to acoustic condition changes. A possible solution is to base the BN features on a channel-robust primary feature. In this paper, we propose to derive the BN feature based on Gammatone frequency cepstral coefficients (GFCCs). The GFCC feature has shown nice robustness against acoustic change, due to its capability of simulating the auditory system of humans. The idea is to integrate the advantage of the GFCC feature in acoustic robustness and the advantage of the BN feature in signal representation, so that the BN feature can be improved in the condition of mismatched training/test channels. This is particularly useful for small-scale tasks for which the training data are often limited. The experiments are conducted on the WSJCAM0 database, where the test utterances are mixed with noises at various SNR levels to simulate the channel change. The results confirm that the GFCC-based BN feature is much more robust than the BN features based on the MFCC and the PLP. Furthermore, the primary GFCC feature and the GFCC-based BN feature can be concatenated, leading to a more robust combined feature which provides considerable performance gains in all the tested noise conditions. Jun Qi 0002, Dong Wang 0013, Javier Tejedor |
INTERSPEECH | 1 |
| 2013 | Auditory features based on Gammatone filters for robust speech recognitionabstractA major challenge for automatic speech recognition (ASR) relates to significant performance reduction in noisy environments. Recent research has shown that auditory features based on Gammatone filters are promising to improve robustness of ASR systems against noise, though the research is far from extensive and generalizability of the new features is unknown. This paper presents our implementation of the Gamma-tone filter-based feature and the experimental results on Mandarin speech data. By some thorough designs, we obtained significant performance gains with the new feature in various noise conditions when compared with the widely used MFCC and PLP features. A particular novelty of our implementation is that the filter design is purely in the time domain. This means that the channel signals are obtained with a set of Gammatone filters applied directly on the speech signals in time domain, which is totally different from the commonly adopted frequency-domain design that first converts signals to spectra and then applies the filter banks upon them. The time-domain implementation on the one hand avoids the approximation introduced by short-time spectral analysis and hence is more precise; and on the other hand, it avoids the complex spectral computation and hence simplifies hardware realization. Jun Qi 0002, Dong Wang 0013, Runsheng Liu |
ISCAS | 1 |