EDBT 2026 Demo / reviewers in the wild / expert
Zheng-Hua Tan
dblp:39/4898
· DBLP profile ↗
143ranked-venue papers
13as first author
41since 2021 · last 2025
0000-0001-6856-8928ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 88 · 10 first-author · 24 since 2021Artificial intelligence and machine learning · 87 · 8 first-author · 24 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-authorSystems, architecture and hardware · 3Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Detecting and Defending Against Adversarial Attacks on Automatic Speech Recognition via Diffusion ModelsabstractAutomatic speech recognition (ASR) systems are known to be vulnerable to adversarial attacks. This paper addresses detection and defence against targeted white-box attacks on speech signals for ASR systems. While existing work has utilised diffusion models (DMs) to purify adversarial examples, achieving state-of-the-art results in keyword spotting tasks, their effectiveness for more complex tasks such as sentence-level ASR remains unexplored. Additionally, the impact of the number of forward diffusion steps on performance is not well understood. In this paper, we systematically investigate the use of DMs for defending against adversarial attacks on sentences and examine the effect of varying forward diffusion steps. Through comprehensive experiments on the Mozilla Common Voice dataset, we demonstrate that two forward diffusion steps can completely defend against adversarial attacks on sentences. Moreover, we introduce a novel, training-free approach for detecting adversarial attacks by leveraging a pre-trained DM. Our experimental results show that this method can detect adversarial attacks with high accuracy. Nikolai Lund Kühne, Astrid H. F. Kitchena, Marie S. Jensen, Mikkel S. L. Brøndt, Martin Gonzalez, Christophe Biscio, Zheng-Hua Tan |
ICASSP | 7 |
| 2025 | Deep Feedback Cancellation for Hearing Aids with Improved System Stability and Sound QualityabstractAcoustic feedback cancellation is an important task in audio processing systems, aiming to mitigate the effects of feedback loops on system stability and sound quality. State-of-the-art methods rely on adaptive filtering algorithms and face challenges in balancing between rapid convergence and low steady-state error. In this work, we introduce a novel approach inspired by traditional adaptive filtering and deep learning techniques to achieve a significantly faster convergence and lower steady-state errors at the same time. Our proposed system, termed Deep Feedback Cancellation (DFC), leverages deep neural networks to predict the impulse response of the feedback path directly. Hence, it replaces traditional gradient based adaptive estimation of the impulse responses. Experimental evaluation, in a hearing aid setting, conducted on real-world data demonstrates the superiority of DFC over traditional methods. Specifically, in a practically very important situation, where the feedback path undergoes rapid changes, the proposed DFC achieves increased convergence rate by a factor of 30, while decreasing the steady-state error by 2 dB. Generally, DFC leads to very significant improvements over traditional methods. These improvements are confirmed by objective evaluations and subjective listening tests. Our findings suggest that DFC presents a promising alternative for acoustic feedback cancellation in hearing aid applications. Eleftheria Lydaki, Zheng-Hua Tan, Jesper Jensen 0001, Meng Guo 0001 |
ICASSP | 2 |
| 2025 | Optimal Sensor Scheduling and Selection for Continuous-Discrete Kalman Filtering with Auxiliary DynamicsabstractWe study the Continuous-Discrete Kalman Filter (CD-KF) for State-Space Models (SSMs) where continuous-time dynamics are observed via multiple sensors with discrete, irregularly timed measurements. Our focus extends to scenarios in which the measurement process is coupled with the states of an auxiliary SSM. For instance, higher measurement rates may increase energy consumption or heat generation, while a sensor’s accuracy can depend on its own spatial trajectory or that of the measured target. Each sensor thus carries distinct costs and constraints associated with its measurement rate and additional constraints and costs on the auxiliary state. We model measurement occurrences as independent Poisson processes with sensor-specific rates and derive an upper bound on the mean posterior covariance matrix of the CD-KF along the mean auxiliary state. The bound is continuously differentiable with respect to the measurement rates, which enables efficient gradient-based optimization. Exploiting this bound, we propose a finite-horizon optimal control framework to optimize measurement rates and auxiliary-state dynamics jointly. We further introduce a deterministic method for scheduling measurement times from the optimized rates. Empirical results in state-space filtering and dynamic temporal Gaussian process regression demonstrate that our approach achieves improved trade-offs between resource usage and estimation accuracy. Mohamad Al Ahdab, John Leth, Zheng-Hua Tan |
ICML | 3 |
| 2025 | xLSTM-SENet: xLSTM for Single-Channel Speech EnhancementabstractWhile attention-based architectures, such as Conformers, excel in speech enhancement, they face challenges such as scalability with respect to input sequence length. In contrast, the recently proposed Extended Long Short-Term Memory (xLSTM) architecture offers linear scalability. However, xLSTM-based models remain unexplored for speech enhancement. This paper introduces xLSTM-SENet, the first xLSTM-based single-channel speech enhancement system. A direct comparative analysis reveals that xLSTM-and notably, even LSTM-can match or outperform state-of-the-art Mamba- and Conformer-based systems across various model sizes in speech enhancement on the VoiceBank+Demand dataset. Through ablation studies, we identify key architectural design choices such as exponential gating and bidirectionality contributing to its effectiveness. Our best xLSTM-based model, xLSTM-SENet2, outperforms state-of-the-art Mamba- and Conformer-based systems of similar complexity on the Voicebank+DEMAND dataset. Nikolai Lund Kühne, Jan Østergaard, Jesper Jensen 0001, Zheng-Hua Tan |
INTERSPEECH | 4 |
| 2025 | Analysis and Extension of a Near-End Listening Enhancement Method Based on Long-Term Fractile Noise StatisticsabstractThis paper addresses the problem of near-end listening enhancement (NELE), where a clean speech signal is modified prior to playback and under an energy constraint to improve intelligibility in noise. We analyze a recently proposed NELE method, optimized using a Speech Intelligibility Index that has been modified to incorporate temporal aspects of the noise via long-term fractile noise statistics. Specifically, we explain the energy allocation strategy adopted by the algorithm, and show that, in contrast to many existing methods, the spectral energy distribution of the modified speech is a function of that of the background noise, but not that of the input speech. Our simulation experiments show that this simple method outperforms well-established spectral shaping NELE methods. In addition, we extend the algorithm by appending an off-the-shelf dynamic range compressor, and show that it performs generally better than state-of-the-art methods for NELE. Filippo Villani, Wai-Yip Chan, Zheng-Hua Tan, Jan Østergaard, Jesper Jensen 0001 |
INTERSPEECH | 3 |
| 2025 | AxLSTMs: learning self-supervised audio representations with xLSTMsabstractWhile the transformer has emerged as the eminent neural architecture, several independent lines of research have emerged to address its limitations. Recurrent neural approaches have observed a lot of renewed interest, including the extended long short-term memory (xLSTM) architecture, which reinvigorates the original LSTM. However, while xLSTMs have shown competitive performance compared to the transformer, their viability for learning self-supervised general-purpose audio representations has not been evaluated. This work proposes Audio xLSTM (AxLSTM), an approach for learning audio representations from masked spectrogram patches in a self-supervised setting. Pretrained on the AudioSet dataset, the proposed AxLSTM models outperform comparable self-supervised audio spectrogram transformer (SSAST) baselines by up to 25% in relative performance across a set of ten diverse downstream tasks while having up to 45% fewer parameters. Sarthak Yadav, Sergios Theodoridis, Zheng-Hua Tan |
INTERSPEECH | 3 |
| 2025 | A survey of deep learning for complex speech spectrogramsabstractRecent advancements in deep learning have significantly impacted the field of speech signal processing, particularly in the analysis and manipulation of complex spectrograms. This survey provides a comprehensive overview of the state-of-the-art techniques leveraging deep neural networks for processing complex spectrograms, which encapsulate both magnitude and phase information. We begin by introducing complex spectrograms and their associated features for various speech processing tasks. Next, we examine the key components and architectures of complex-valued neural networks, which are specifically designed to handle complex-valued data and have been applied to complex spectrogram processing. As recent studies have primarily focused on applying real-valued neural networks to complex spectrograms, we revisit these approaches and their architectural designs. We then discuss various training strategies and loss functions tailored for training neural networks to process and model complex spectrograms. The survey further examines key applications, including phase retrieval, speech enhancement, and speaker separation, where deep learning has achieved significant progress by leveraging complex spectrograms or their derived feature representations. Additionally, we examine the intersection of complex spectrograms with generative models. This survey aims to serve as a valuable resource for researchers and practitioners in the field of speech signal processing, deep learning and related fields. Yuying Xie 0002, Zheng-Hua Tan |
Speech Commun. | 2 |
| 2024 | PAC-Bayes Generalisation Bounds for Dynamical Systems including Stable RNNsabstractIn this paper, we derive a PAC-Bayes bound on the generalisation gap, in a supervised time-series setting for a special class of discrete-time non-linear dynamical systems. This class includes stable recurrent neural networks (RNN), and the motivation for this work was its application to RNNs. In order to achieve the results, we impose some stability constraints, on the allowed models. Here, stability is understood in the sense of dynamical systems. For RNNs, these stability conditions can be expressed in terms of conditions on the weights. We assume the processes involved are essentially bounded and the loss functions are Lipschitz. The proposed bound on the generalisation gap depends on the mixing coefficient of the data distribution, and the essential supremum of the data. Furthermore, the bound converges to zero as the dataset size increases. In this paper, we 1) formalize the learning problem, 2) derive a PAC-Bayesian error bound for such systems, 3) discuss various consequences of this error bound, and 4) show an illustrative example, with discussions on computing the proposed bound. Unlike other available bounds the derived bound holds for non i.i.d. data (time-series) and it does not grow with the number of steps of the RNN. Deividas Eringis, John Leth, Zheng-Hua Tan, Rafael Wisniewski, Mihály Petreczky |
AAAI | 3 |
| 2024 | Self-Supervised Pretraining for Robust Personalized Voice Activity Detection in Adverse ConditionsabstractIn this paper, we propose the use of self-supervised pretraining on a large unlabelled data set to improve the performance of a personalized voice activity detection (VAD) model in adverse conditions. We pretrain a long short-term memory (LSTM)-encoder using the autoregressive predictive coding (APC) framework and fine-tune it for personalized VAD. We also propose a denoising variant of APC, with the goal of improving the robustness of personalized VAD. The trained models are systematically evaluated on both clean speech and speech contaminated by various types of noise at different SNR-levels and compared to a purely supervised model. Our experiments show that self-supervised pretraining not only improves performance in clean conditions, but also yields models which are more robust to adverse conditions compared to purely supervised learning. Holger S. Bovbjerg, Jesper Jensen 0001, Jan Østergaard, Zheng-Hua Tan |
ICASSP | 4 |
| 2024 | Diffusion-Based Speech Enhancement in Matched and Mismatched Conditions Using a Heun-Based SamplerabstractDiffusion models are a new class of generative models that have recently been applied to speech enhancement successfully. Previous works have demonstrated their superior performance in mismatched conditions compared to state-of-the art discriminative models. However, this was investigated with a single database for training and another one for testing, which makes the results highly dependent on the particular databases. Moreover, recent developments from the image generation literature remain largely unexplored for speech enhancement. These include several design aspects of diffusion models, such as the noise schedule or the reverse sampler. In this work, we systematically assess the generalization performance of a diffusion-based speech enhancement model by using multiple speech, noise and binaural room impulse response (BRIR) databases to simulate mismatched acoustic conditions. We also experiment with a noise schedule and a sampler that have not been applied to speech enhancement before. We show that the proposed system substantially benefits from using multiple databases for training, and achieves superior performance compared to state-of-the-art discriminative models in both matched and mismatched conditions. We also show that a Heun-based sampler achieves superior performance at a smaller computational cost compared to a sampler commonly used for speech enhancement. Philippe Gonzalez, Zheng-Hua Tan, Jan Østergaard, Jesper Jensen 0001, Tommy S. Alstrøm, Tobias May |
ICASSP | 2 |
| 2024 | Masked Autoencoders with Multi-Window Local-Global Attention Are Better Audio LearnersabstractIn this work, we propose a Multi-Window Masked Autoencoder (MW-MAE) fitted with a novel Multi-Window Multi-Head Attention (MW-MHA) module that facilitates the modelling of local-global interactions in every decoder transformer block through attention heads of several distinct local and global windows. Empirical results on ten downstream audio tasks show that MW-MAEs consistently outperform standard MAEs in overall performance and learn better general-purpose audio representations, along with demonstrating considerably better scaling characteristics. Investigating attention distances and entropies reveals that MW-MAE encoders learn heads with broader local and global attention. Analyzing attention head feature representations through Projection Weighted Canonical Correlation Analysis (PWCCA) shows that attention heads with the same window sizes across the decoder layers of the MW-MAE learn correlated feature representations which enables each block to independently capture local and global information, leading to a decoupled decoder feature hierarchy. Sarthak Yadav, Sergios Theodoridis, Lars Kai Hansen, Zheng-Hua Tan |
ICLR | 4 |
| 2024 | PAC-Bayesian Error Bound, via Rényi Divergence, for a Class of Linear Time-Invariant State-Space ModelsabstractIn this paper we derive a PAC-Bayesian error bound for a class of stochastic dynamical systems with inputs, namely, for linear time-invariant stochastic state-space models (stochastic LTI systems for short). This class of systems is widely used in control engineering and econometrics, in particular, they represent a special case of recurrent neural networks. In this paper we 1) formalize the learning problem for stochastic LTI systems with inputs, 2) derive a PAC-Bayesian error bound for such systems, and 3) discuss various consequences of this error bound. Deividas Eringis, John Leth, Zheng-Hua Tan, Rafael Wisniewski, Mihály Petreczky |
ICML | 3 |
| 2024 | Complex Recurrent Variational Autoencoder for Speech Resynthesis and EnhancementabstractAiming at learning a probabilistic distribution over data, generative models have been actively studied with broad applications. This paper proposes a complex recurrent variational autoencoder (VAE) framework, for modeling time series data, particularly speech signals. First, to account for the temporal structure of speech signals, we introduce complex-valued recurrent neural network in the framework. Then, inspired by recent advancements in speech enhancement and separation, the reconstruction loss in the proposed model is L1-based loss, considering penalty on both complex and magnitude spectrograms. To exemplify the use of the complex generative model, we choose speech resynthesis first and then enhancement as the specific application in this paper. Experiments are conducted on the VCTK, TIMIT, and VoiceBank+DEMAND datasets. The results show that the proposed method can resynthesize complex spectrogram well, and offers improvements on objective metrics in speech intelligibility and signal quality for enhancement. Yuying Xie 0002, Thomas Arildsen, Zheng-Hua Tan |
IJCNN | 3 |
| 2024 | Audio Mamba: Selective State Spaces for Self-Supervised Audio RepresentationsabstractDespite its widespread adoption as the prominent neural architecture, the Transformer has spurred several independent lines of work to address its limitations. One such approach is selective state space models, which have demonstrated promising results for language modelling. However, their feasibility for learning self-supervised, general-purpose audio representations is yet to be investigated. This work proposes Audio Mamba, a selective state space model for learning general-purpose audio representations from randomly masked spectrogram patches through self-supervision. Empirical results on ten diverse audio recognition downstream tasks show that the proposed models, pretrained on the AudioSet dataset, consistently outperform comparable self-supervised audio spectrogram transformer (SSAST) baselines by a considerable margin and demonstrate better performance in dataset size, sequence length and model size comparisons. Sarthak Yadav, Zheng-Hua Tan |
INTERSPEECH | 2 |
| 2024 | The Effect of Training Dataset Size on Discriminative and Diffusion-Based Speech Enhancement SystemsabstractThe performance of deep neural network-based speech enhancement systems typically increases with the training dataset size. However, studies that investigated the effect of training dataset size on speech enhancement performance did not consider recent approaches, such as diffusion-based generative models. Diffusion models are typically trained with massive datasets for image generation tasks, but whether this is also required for speech enhancement is unknown. Moreover, studies that investigated the effect of training dataset size did not control for the data diversity. It is thus unclear whether the performance improvement was due to the increased dataset size or diversity. Therefore, we systematically investigate the effect of training dataset size on the performance of popular state-of-the-art discriminative and diffusion-based speech enhancement systems in matched conditions. We control for the data diversity by using a fixed set of speech utterances, noise segments and binaural room impulse responses to generate datasets of different sizes. We find that the diffusion-based systems perform the best relative to the discriminative systems in terms of objective metrics with datasets of 10 h or less. However, their objective metrics performance does not improve when increasing the training dataset size as much as the discriminative systems, and they are outperformed by the discriminative systems with datasets of 100 h or more. Philippe Gonzalez, Zheng-Hua Tan, Jan Østergaard, Jesper Jensen 0001, Tommy S. Alstrøm, Tobias May |
IEEE Signal Process. Lett. | 2 |
| 2024 | Generating Accurate and Diverse Audio Captions Through Variational Autoencoder FrameworkabstractGenerating both diverse and accurate descriptions is an essential goal in the audio captioning task. Traditional methods mainly focus on improving the accuracy of the generated captions but ignore their diversity. In contrast, recent methods have considered generating diverse captions for a given audio clip, but with the potential trade-off in caption accuracy. In this work, we propose a new diverse audio captioning method based on a variational autoencoder structure, dubbed AC-VAE, aiming to achieve a better trade-off between the diversity and accuracy of the generated captions. To improve diversity, AC-VAE learns the latent word distribution at each location based on contextual information. To uphold accuracy, AC-VAE incorporates an autoregressive prior module and a global constraint module, which enable precise modeling of word distribution and encourage semantic consistency of captions at the sentence level. We evaluate the proposed AC-VAE on the Clotho dataset. Experimental results show that AC-VAE achieves a better trade-off between diversity and accuracy compared to the state-of-the-art methods. The code is publicly available at https://github.com/XinMing0411/AC-VAE Yiming Zhang 0025, Ruoyi Du, Zheng-Hua Tan, Wenwu Wang 0001, Zhanyu Ma |
IEEE Signal Process. Lett. | 3 |
| 2024 | Investigating the Design Space of Diffusion Models for Speech EnhancementabstractDiffusion models are a new class of generative models that have shown outstanding performance in image generation literature. As a consequence, studies have attempted to apply diffusion models to other tasks, such as speech enhancement. A popular approach in adapting diffusion models to speech enhancement consists in modelling a progressive transformation between the clean and noisy speech signals. However, one popular diffusion model framework previously laid in image generation literature did not account for such a transformation towards the system input, which prevents from relating the existing diffusion-based speech enhancement systems with the aforementioned diffusion model framework. To address this, we extend this framework to account for the progressive transformation between the clean and noisy speech signals. This allows us to apply recent developments from image generation literature, and to systematically investigate design aspects of diffusion models that remain largely unexplored for speech enhancement, such as the neural network preconditioning, the training loss weighting, the stochastic differential equation (SDE), or the amount of stochasticity injected in the reverse process. We show that the performance of previous diffusion-based speech enhancement systems cannot be attributed to the progressive transformation between the clean and noisy speech signals. Moreover, we show that a proper choice of preconditioning, training loss weighting, SDE and sampler allows to outperform a popular diffusion-based speech enhancement system while using fewer sampling steps, thus reducing the computational cost by a factor of four. Philippe Gonzalez, Zheng-Hua Tan, Jan Østergaard, Jesper Jensen 0001, Tommy S. Alstrøm, Tobias May |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2024 | How to Train Your Ears: Auditory-Model Emulation for Large-Dynamic-Range Inputs and Mild-to-Severe Hearing LossesabstractAdvanced auditory models are useful in designing signal-processing algorithms for hearing-loss compensation or speech enhancement. Such auditory models provide rich and detailed descriptions of the auditory pathway, and might allow for individualization of signal-processing strategies, based on physiological measurements. However, these auditory models are often computationally demanding, requiring significant time to compute. To address this issue, previous studies have explored the use of deep neural networks to emulate auditory models and reduce inference time. While these deep neural networks offer impressive efficiency gains in terms of computational time, they may suffer from uneven emulation performance as a function of auditory-model frequency-channels and input sound pressure level, making them unsuitable for many tasks. In this study, we demonstrate that the conventional machine-learning optimization objective used in existing state-of-the-art methods is the primary source of this limitation. Specifically, the optimization objective fails to account for the frequency- and level-dependencies of the auditory model, caused by a large input dynamic range and different types of hearing losses emulated by the auditory model. To overcome this limitation, we propose a new optimization objective that explicitly embeds the frequency- and level-dependencies of the auditory model. Our results show that this new optimization objective significantly improves the emulation performance of deep neural networks across relevant input sound levels and auditory-model frequency channels, without increasing the computational load during inference. Addressing these limitations is essential for advancing the application of auditory models in signal-processing tasks, ensuring their efficacy in diverse scenarios. Peter Leer, Jesper Jensen 0001, Zheng-Hua Tan, Jan Østergaard, Lars Bramslow |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2024 | Data-Driven Non-Intrusive Speech Intelligibility Prediction Using Speech Presence ProbabilityabstractTime consuming Speech Intelligibility (SI) listening tests with human subjects can be replaced by algorithmic SI predictors. In recent years, data-driven SI predictors have been showing promising results. A major limiting factor in the advancement of data-driven SI prediction is that there is a scarcity of SI listening test data available to train the data-driven methods. In this article we propose a data-driven SI predictor that does not require access to an underlying noise-free reference signal, i.e.,non-intrusive, and which does not require listening test data for training. Instead, the proposed method exploits a hypothesized link between SI and Speech Presence Probability (SPP). We show that a neural network can be trained on easily obtainable speech in additive noise data to estimate SPP, and that a simple post-processing stage can be applied in order to map the estimated SPP to SI predictions with high accuracy. The proposed method is evaluated and compared to other state-of-the art non-intrusive SI predictors, and achieves the highest performance even in the presence of processed noisy speech, which the SPP estimator has not been trained on. Mathias Bach Pedersen, Søren Holdt Jensen, Zheng-Hua Tan, Jesper Jensen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Filterbank Learning for Noise-Robust Small-Footprint Keyword SpottingabstractIn the context of keyword spotting (KWS), the replacement of handcrafted speech features by learnable features has not yielded superior KWS performance. In this study, we demonstrate that filterbank learning outperforms handcrafted speech features for KWS whenever the number of filterbank channels is severely decreased. Reducing the number of channels might yield certain KWS performance drop, but also a substantial energy consumption reduction, which is key when deploying common always-on KWS on low-resource devices. Experimental results on a noisy version of the Google Speech Commands Dataset show that filterbank learning adapts to noise characteristics to provide a higher degree of robustness to noise, especially when dropout is integrated. Thus, switching from typically used 40-channel log-Mel features to 8channel learned features leads to a relative KWS accuracy loss of only 3.5% while simultaneously achieving a 6.3× energy consumption reduction. Iván López-Espejo, Ram C. M. C. Shekar, Zheng-Hua Tan, Jesper Jensen 0001, John H. L. Hansen |
ICASSP | 3 |
| 2023 | Radio Sensing with Large Intelligent Surface for 6GabstractThis paper leverages the potential of Large Intelligent Surfaces (LIS) for radio sensing in 6G wireless networks. By taking advantage of arbitrary communication signals occurring in the scenario, we apply direct processing to the output signal from the LIS to obtain a radio map that describes the physical presence of passive devices (scatterers, humans) which act as virtual sources due to the communication signal reflections. We then assess the usage of machine learning and computer vision methods including clustering, template matching and component labeling to extract meaningful information from these radio maps. As an exemplary use case, we evaluate this method for passive multi-human detection in an indoor setting. The results show that the presented method has high application potential as we are able to detect around 98% of humans passively even in quite unfavorable Signal-to-Noise Ratio (SNR) conditions. Cristian J. Vaca-Rubio, Pablo Ramirez-Espinosa, Kimmo Kansanen, Zheng-Hua Tan, Elisabeth de Carvalho |
ICASSP | 4 |
| 2023 | Speech inpainting: Context-based speech synthesis guided by videoabstractAudio and visual modalities are inherently connected in speech signals: lip movements and facial expressions are correlated with speech sounds. This motivates studies that incorporate the visual modality to enhance an acoustic speech signal or even restore missing audio information. Specifically, this paper focuses on the problem of audio-visual speech inpainting, which is the task of synthesizing the speech in a corrupted audio segment in a way that it is consistent with the corresponding visual content and the uncorrupted audio context. We present an audio-visual transformer-based deep learning model that leverages visual cues that provide information about the content of the corrupted audio. It outperforms the previous state-of-the-art audio-visual model and audio-only baselines. We also show how visual features extracted with AV-HuBERT, a large audiovisual transformer for speech recognition, are suitable for synthesizing speech. Juan F. Montesinos, Daniel Michelsanti, Gloria Haro, Zheng-Hua Tan, Jesper Jensen 0001 |
INTERSPEECH | 4 |
| 2023 | On the deficiency of intelligibility metrics as proxies for subjective intelligibilityabstractA recent trend in deep neural network (DNN)-based speech enhancement consists of using intelligibility and quality metrics as loss functions for model training with the aim of achieving high subjective speech intelligibility and perceptual quality in real-life conditions. In this study, we analyze a variety of loss functions, including some based on state-of-the-art intelligibility and quality metrics, to train an end-to-end speech enhancement system based on a fully convolutional neural network. The loss functions include perceptual metric for speech quality evaluation (PMSQE), scale-invariant signal-to-distortion ratio (SI-SDR), SI-SDR integrating speech pre-emphasis, short-time objective intelligibility (STOI), extended STOI (ESTOI), spectro-temporal glimpsing index (STGI), and a composite loss function combining STGI and SI-SDR. While DNNs trained with these loss functions produce notable speech intelligibility (and quality) gains according to pertinent objective metrics, we conduct a subjective intelligibility test that contradicts this result, showing no intelligibility improvement. From the results of this study, our conclusion is twofold: (1) subjective intelligibility evaluation is currently not replaceable by objective intelligibility evaluation, and (2) both the development of meaningful intelligibility metrics and DNN-based speech enhancement systems that can consistently improve the intelligibility of noisy speech for human listening remain open problems. Iván López-Espejo, Amin Edraki, Wai-Yip Chan, Zheng-Hua Tan, Jesper Jensen 0001 |
Speech Commun. | 4 |
| 2023 | Minimum Processing Near-End Listening EnhancementabstractThe intelligibility and quality of speech from a mobile phone or public announcement system are often affected by background noise in the listening environment. By pre-processing the speech signal it is possible to improve the speech intelligibility and quality — this is known as near-end listening enhancement (NLE). Although, existing NLE techniques are able to greatly increase intelligibility in harsh noise environments, in favorable noise conditions the intelligibility of speech reaches a ceiling where it cannot be further enhanced. Actually, the focus of existing methods solely on improving the intelligibility causes unnecessary processing of the speech signal and leads to speech distortions and quality degradations. In this article, we provide a new rationale for NLE, where the target speech is minimally processed in terms of a processing penalty, provided that a certain performance constraint, e.g., intelligibility, is satisfied. We present a closed-form solution for the case where the performance criterion is an intelligibility estimator based on the approximated speech intelligibility index and the processing penalty is the mean-square error between the processed and the clean speech. This produces an NLE method that adapts to changing noise conditions via a simple gain rule by limiting the processing to the minimum necessary to achieve a desired intelligibility, while at the same time focusing on quality in favorable noise situations by minimizing the amount of speech distortions. Through simulation studies, we show the proposed method attains speech quality on par or better than existing methods in both objective measurements and subjective listening tests, whilst still sustaining objective speech intelligibility performance on par with existing methods. Andreas Jonas Fuglsig, Jesper Jensen 0001, Zheng-Hua Tan, Lars Søndergaard Bertelsen, Jens Christian Lindof, Jan Østergaard |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | ACTUAL: Audio Captioning With Caption Feature Space RegularizationabstractAudio captioning aims at describing the content of audio clips with human language. Due to the ambiguity of audio content, different people may perceive the same audio clip differently, resulting in caption disparities (i.e., the same audio clip may be described by several captions with diverse semantics). In the literature, the one-to-many strategy is often employed to train the audio captioning models, where a related caption is randomly selected as the optimization target for each audio clip at each training iteration. However, we observe that this can lead to significant variations during the optimization process and adversely affect the performance of the model. In this paper, we address this issue by proposing an audio captioning method, named ACTUAL (Audio Captioning with capTion featUre spAce reguLarization). ACTUAL involves a two-stage training process: (i) in the first stage, we use contrastive learning to construct a proxy feature space where the similarities between captions at the audio level are explored, and (ii) in the second stage, the proxy feature space is utilized as additional supervision to improve the optimization of the model in a more stable direction. We conduct extensive experiments to demonstrate the effectiveness of the proposed ACTUAL method. The results show that proxy caption embedding can significantly improve the performance of the baseline model and the proposed ACTUAL method offers competitive performance on two datasets compared to state-of-the-art methods. The code is publicly available athttps://github.com/PRIS-CV/Caption-Feature-Space-Regularization. Yiming Zhang 0025, Hong Yu 0006, Ruoyi Du, Zheng-Hua Tan, Wenwu Wang 0001, Zhanyu Ma |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | On the Comparisons of Decorrelation Approaches for Non-Gaussian Neutral Vector Variablesabstract-norm equals one. In addition, its neutral properties make it significantly different from the commonly studied vector variables (e.g., the Gaussian vector variables). Due to the aforementioned properties, the conventionally applied linear transformation approaches [e.g., principal component analysis (PCA) and independent component analysis (ICA)] are not suitable for neutral vector variables, as PCA cannot transform a neutral vector variable, which is highly negatively correlated, into a set of mutually independent scalar variables and ICA cannot preserve the bounded property after transformation. In recent work, we proposed an efficient nonlinear transformation approach, i.e., the parallel nonlinear transformation (PNT), for decorrelating neutral vector variables. In this article, we extensively compare PNT with PCA and ICA through both theoretical analysis and experimental evaluations. The results of our investigations demonstrate the superiority of PNT for decorrelating the neutral vector variables. Zhanyu Ma, Xiaoou Lu, Jiyang Xie 0001, Zhen Yang 0004, Jing-Hao Xue, Zheng-Hua Tan, Bo Xiao 0006, Jun Guo 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2022 | Joint Far- and Near-End Speech Intelligibility Enhancement Based on the Approximated Speech Intelligibility IndexabstractThis paper considers speech enhancement of signals picked up in one noisy environment which must be presented to a listener in another noisy environment. Recently, it has been shown that an optimal solution to this problem requires the consideration of the noise sources in both environments jointly. However, the existing optimal mutual information based method requires a complicated system model that includes natural speech variations, and relies on approximations and assumptions of the underlying signal distributions. In this paper, we propose to use a simpler signal model and optimize speech intelligibility based on the Approximated Speech Intelligibility Index (ASII). We derive a closed-form solution to the joint far- and near-end speech enhancement problem that is independent of the marginal distribution of signal coefficients, and that achieves similar performance to existing work. In addition, we do not need to model or optimize for natural speech variations. Andreas Jonas Fuglsig, Jan Østergaard, Jesper Jensen 0001, Lars Søndergaard Bertelsen, Peter Mariager, Zheng-Hua Tan |
ICASSP | 6 |
| 2022 | Summary on the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand ChallengeabstractThe ICASSP 2022 Multi-channel Multi-party Meeting Transcription Grand Challenge (M2MeT) focuses on one of the most valuable and the most challenging scenarios of speech technologies. The M2MeT challenge has particularly set up two tracks, speaker diarization (track 1) and multi-speaker automatic speech recognition (ASR) (track 2). Along with the challenge, we released 120 hours of real-recorded Mandarin meeting speech data with manual annotation, including far-field data collected by 8-channel micro-phone array as well as near-field data collected by each participants’ headset microphone. We briefly describe the released dataset, track setups, baselines and summarize the challenge results and major techniques used in the submissions. Fan Yu 0002, Shiliang Zhang, Yihui Fu, Zhihao Du, Weilong Huang, Lei Xie 0001, Zheng-Hua Tan, DeLiang Wang, Yanmin Qian, Kong-Aik Lee, Zhijie Yan, Bin Ma 0001, Hui Bu |
ICASSP | 9 |
| 2022 | Adversarial Multi-Task Deep Learning for Noise-Robust Voice Activity Detection with Low Algorithmic DelayabstractVoice Activity Detection (VAD) is an important pre-processing step in a wide variety of speech processing systems. VAD should in a practical application be able to detect speech in both noisy and noise-free environments, while not introducing significant latency. In this work we propose using an adversarial multi-task learning method when training a supervised VAD. The method has been applied to the state-of-the-art VAD Waveform-based Voice Activity Detection. Additionally the performance of the VAD is investigated under different algorithmic delays, which is an important factor in latency. Introducing adversarial multi-task learning to the model is observed to increase performance in terms of Area Under Curve (AUC), particularly in noisy environments, while the performance is not degraded at higher SNR levels. The adversarial multi-task learning is only applied in the training phase and thus introduces no additional cost in testing. Furthermore the correlation between performance and algorithmic delays is investigated, and it is observed that the VAD performance degradation is only moderate when lowering the algorithmic delay from 398 ms to 23 ms. Index Terms: Voice Activity Detection, adversarial multi-task learning, algorithmic delay, deep learning, noise robustness. Claus M. Larsen, Peter Koch 0001, Zheng-Hua Tan |
INTERSPEECH | 3 |
| 2022 | AoI and Throughput Optimization for Hybrid Traffic in Cellular Uplink Using Reinforcement LearningabstractThe fast growth of time-sensitive applications calls for the optimization of radio access network (RAN) scheduling. We consider the problem of RAN scheduling of a mix of periodic and burst traffic and design a reinforcement learning method for the age of information and throughput optimization. The periodic traffic is generated with a fixed frequency and the burst traffic is generated by the Poisson Pareto Burst Process. We firstly formulate the scheduling problem as a non-linear integer programming problem. Then, we focus on the reinforcement learning method modeling and solve it via the Proximal Policy Optimization algorithm. Our evaluations show that the suggested reinforcement algorithm outperforms the classical algorithms without any prior knowledge of the arriving traffic. Chien-Cheng Wu, Zheng-Hua Tan, Cedomir Stefanovic |
VTC Spring | 2 |
| 2022 | Advanced Dropout: A Model-Free Methodology for Bayesian Dropout OptimizationabstractDue to lack of data, overfitting ubiquitously exists in real-world applications of deep neural networks (DNNs). We propose advanced dropout, a model-free methodology, to mitigate overfitting and improve the performance of DNNs. The advanced dropout technique applies a model-free and easily implemented distribution with parametric prior, and adaptively adjusts dropout rate. Specifically, the distribution parameters are optimized by stochastic gradient variational Bayes in order to carry out an end-to-end training. We evaluate the effectiveness of the advanced dropout against nine dropout techniques on seven computer vision datasets (five small-scale datasets and two large-scale datasets) with various base models. The advanced dropout outperforms all the referred techniques on all the datasets. We further compare the effectiveness ratios and find that advanced dropout achieves the highest one on most cases. Next, we conduct a set of analysis of dropout rate characteristics, including convergence of the adaptive dropout rate, the learned distributions of dropout masks, and a comparison with dropout rate generation without an explicit distribution. In addition, the ability of overfitting prevention is evaluated and confirmed. Finally, we extend the application of the advanced dropout to uncertainty inference, network pruning, text classification, and regression. The proposed advanced dropout is also superior to the corresponding referred methods. Codes are available at https://github.com/PRIS-CV/AdvancedDropout. Jiyang Xie 0001, Zhanyu Ma, Jianjun Lei 0001, Guoqiang Zhang 0003, Jing-Hao Xue, Zheng-Hua Tan, Jun Guo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Multichannel Speech Enhancement With Own Voice-Based Interfering Speech Suppression for Hearing Assistive DevicesabstractEnhancementof a desired speech signal in the presence of competing or interfering speech remains an unsolved problem, as it can be hard to determine which of the speech signals is the one of interest. In this paper, we propose a multichannel noise reduction algorithm which uses the presence of the user’s own voice signal, e.g. during conversations with the target speaker, as an asset to efficiently identify interfering speech and noise. Specifically, following the typical speech pattern in natural conversations, the presence of an own voice may indicate the absence of the target speech, hence undesired speech and noise can be identified and estimated during own voice presence. In contrast to conventional noise reduction systems, the proposed noise reduction systems use the user’s own voice to identify interfering speech that otherwise could be confused with the target speech. We demonstrate the performance of the proposed noise reduction systems in a comparison against state-of-the-art noise reduction systems in terms of beamforming performance for hearing assistive devices. The results show that the proposed beamforming scheme in particular outperforms state-of-the-art methods in terms of ESTOI and PESQ in situations with a target speaker and a strong interfering speaker. Poul Hoang, Jan Mark de Haan, Zheng-Hua Tan, Jesper Jensen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Conferencingspeech Challenge: Towards Far-Field Multi-Channel Speech Enhancement for Video ConferencingabstractThe ConferencingSpeech 2021 challenge is proposed to stimulate research on far-field multi-channel speech enhancement for video conferencing. The challenge consists of two separate tasks: 1) Task 1 is multi-channel speech enhancement with single microphone array and focusing on practical application with real-time requirement and 2) Task 2 is multi-channel speech enhancement with multiple distributed micro-phone arrays, which is a non-real-time track and does not have any constraints so that participants could explore any algorithms to obtain high speech quality. Targeting the real video conferencing room application, the challenge database was recorded from real speakers and all recording facilities were located by following the real setup of conferencing room. In this challenge, we open-sourced the list of open source clean speech and noise datasets, simulation scripts, and a baseline system for participants to develop their own system. The final ranking of the challenge will be decided by the subjective evaluation which is performed using Absolute Category Ratings (ACR) to estimate Mean Opinion Score (MOS), speech MOS (S-MOS), and noise MOS (N-MOS). This paper describes the challenge, tasks, datasets, subjective evaluation, and challenge results. The baseline system which is a complex ratio mask based neural network and its experimental results are also presented. Wei Rao 0002, Yihui Fu, Yanxin Hu, Yvkai Jv, Jiangyu Han, Zhongjie Jiang, Lei Xie 0001, Yannan Wang, Shinji Watanabe 0001, Zheng-Hua Tan, Hui Bu, Shidong Shang |
ASRU | 11 |
| 2021 | Joint Maximum Likelihood Estimation of Power Spectral Densities and Relative Acoustic Transfer Functions for Acoustic BeamformingabstractAcoustic beamforming is crucial for many applications where ex-traction of a target signal from a noisy environment is required. In order to implement practical beamformers, e.g. the multichannel Wiener filter (MWF), estimation of the target and noise power spectral densities (PSDs), and the relative acoustic transfer functions (RATFs) is essential. Several methods, e.g. the so-called covariance whitening (CW) approach, have been proposed for estimating these parameters. However, it seems largely unknown that the CW approach in fact leads to maximum likelihood (ML) estimates of the RATFs. We use historical results to derive joint ML estimates (MLEs) of the RATFs and PSDs in the context of acoustic beam-forming. In addition, based on the MLEs, we propose a basic VAD framework using concentrated likelihood ratios. We use the joint MLEs of the PSDs, RATFs, and the proposed VAD to implement beamformers in a hearing aid application, and compare its performance to competing methods. Simulation results show that the pro-posed scheme can outperform competing methods, in particular in realistic situations where highly accurate prior RATF knowledge is not available or at higher signal-to-noise ratios. Poul Hoang, Zheng-Hua Tan, Jan Mark de Haan, Jesper Jensen 0001 |
ICASSP | 2 |
| 2021 | Audio-Visual Speech Inpainting with Deep LearningabstractIn this paper, we present a deep-learning-based framework for audio-visual speech inpainting, i.e., the task of restoring the missing parts of an acoustic speech signal from reliable audio context and uncorrupted visual information. Recent work focuses solely on audio-only methods and generally aims at inpainting music signals, which show highly different structure than speech. Instead, we inpaint speech signals with gaps ranging from 100 ms to 1600 ms to investigate the contribution that vision can provide for gaps of different duration. We also experiment with a multi-task learning approach where a phone recognition task is learned together with speech inpainting. Results show that the performance of audio-only speech inpainting approaches degrades rapidly when gaps get large, while the proposed audio-visual approach is able to plausibly restore missing information. In addition, we show that multi-task learning is effective, although the largest contribution to performance comes from vision. Giovanni Morrone, Daniel Michelsanti, Zheng-Hua Tan, Jesper Jensen 0001 |
ICASSP | 3 |
| 2021 | UIAI System for Short-Duration Speaker Verification Challenge 2020abstractIn this work, we present the system description of the UIAI entry for the short-duration speaker verification (SdSV) challenge 2020. Our focus is on Task 1 dedicated to text-dependent speaker verification. We investigate different feature extraction and modeling approaches for automatic speaker verification (ASV) and utterance verification (UV). We have also studied different fusion strategies for combining UV and ASV modules. Our primary submission to the challenge is the fusion of seven subsystems which yields a normalized minimum detection cost function (minDCF) of 0.072 and an equal error rate (EER) of 2.14% on the evaluation set. The single system consisting of a pass-phrase identification based model with phone-discriminative bottleneck features gives a normalized minDCF of 0.118 and achieves 19% relative improvement over the state-of-the-art challenge baseline. Md. Sahidullah, Achintya Kumar Sarkar, Ville Vestman, Xuechen Liu 0001, Romain Serizel, Tomi Kinnunen, Zheng-Hua Tan, Emmanuel Vincent 0001 |
SLT | 7 |
| 2021 | Self-segmentation of pass-phrase utterances for deep feature learning in text-dependent speaker verification
Achintya Kumar Sarkar, Zheng-Hua Tan |
Comput. Speech Lang. | 2 |
| 2021 | Deep InterBoost networks for small-sample image classification
Dongliang Chang, Zhanyu Ma, Zheng-Hua Tan, Jing-Hao Xue, Jie Cao 0014, Jun Guo 0002 |
Neurocomputing | 4 |
| 2021 | Vocal Tract Length Perturbation for Text-Dependent Speaker Verification With Autoregressive Prediction CodingabstractIn this letter, we propose a vocal tract length (VTL) perturbation method for text-dependent speaker verification (TD-SV), in which a set of TD-SV systems are trained, one for each VTL factor, and score-level fusion is applied to make a final decision. Next, we explore the bottleneck (BN) feature extracted by training deep neural networks with a self-supervised learning objective, autoregressive predictive coding (APC), for TD-SV and comapre it with the well-studied speaker-discriminant BN feature. The proposed VTL method is then applied to APC and speaker-discriminant BN features. In the end, we combine the VTL perturbation systems trained on MFCC and the two BN features in the score domain. Experiments are performed on the RedDots challenge 2016 database of TD-SV using short utterances with Gaussian mixture model-universal background model and i-vector techniques. Results show the proposed methods significantly outperform the baselines. Achintya Kumar Sarkar, Zheng-Hua Tan |
IEEE Signal Process. Lett. | 2 |
| 2021 | A Novel Loss Function and Training Strategy for Noise-Robust Keyword SpottingabstractThe development of keyword spotting (KWS) systems that are accurate in noisy conditions remains a challenge. Towards this goal, in this paper we propose a novel training strategy relying on multi-condition training for noise-robust KWS. By this strategy, we think of the state-of-the-art KWS models as the composition of a keyword embedding extractor and a linear classifier that are successively trained. To train the keyword embedding extractor, we also propose a new (CN,2+1)-pair loss function extending the concept behind related loss functions like triplet and N-pair losses to reach larger inter-class and smaller intra-class variation. Experimental results on a noisy version of the Google Speech Commands Dataset show that our proposal achieves around 12% KWS accuracy relative improvement with respect to standard end-to-end multi-condition training when speech is distorted by unseen noises. This performance improvement is achieved without increasing the computational complexity of the KWS model. Iván López-Espejo, Zheng-Hua Tan, Jesper Jensen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | An Overview of Deep-Learning-Based Audio-Visual Speech Enhancement and SeparationabstractSpeech enhancement and speech separation are two related tasks, whose purpose is to extract either one or more target speech signals, respectively, from a mixture of sounds generated by several sources. Traditionally, these tasks have been tackled using signal processing and machine learning techniques applied to the available acoustic signals. Since the visual aspect of speech is essentially unaffected by the acoustic environment, visual information from the target speakers, such as lip movements and facial expressions, has also been used for speech enhancement and speech separation systems. In order to efficiently fuse acoustic and visual information, researchers have exploited the flexibility of data-driven approaches, specifically deep learning, achieving strong performance. The ceaseless proposal of a large number of techniques to extract features and fuse multimodal information has highlighted the need for an overview that comprehensively describes and discusses audio-visual speech enhancement and separation based on deep learning. In this paper, we provide a systematic survey of this research topic, focusing on the main elements that characterise the systems in the literature: acoustic features; visual features; deep learning methods; fusion techniques; training targets and objective functions. In addition, we review deep-learning-based methods for speech reconstruction from silent videos and audio-visual sound source separation for non-speech signals, since these methods can be more or less directly applied to audio-visual speech enhancement and separation. Finally, we survey commonly employed audio-visual speech datasets, given their central role in the development of data-driven approaches, and evaluation methods, because they are generally used to compare different systems and determine their performance. Daniel Michelsanti, Zheng-Hua Tan, Shixiong Zhang 0001, Yong Xu 0004, Meng Yu 0003, Dong Yu 0001, Jesper Jensen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Maximum Likelihood Estimation of the Interference-Plus-Noise Cross Power Spectral Density Matrix for Own Voice RetrievalabstractIn headset and hearing aid applications, it is of interest to retrieve the user's own voice in a noisy environment, e.g. for telephony applications. To do so, the cross-power spectral density (CPSD) of the interference-plus-noise is required. In this paper, a novel maximum likelihood (ML) estimator of the interference-plus-noise CPSD matrix is proposed. The proposed method is able to estimate the interference-plus-noise CPSD matrix, even during signal regions with own voice activity. The method uses a novel procedure for estimating the interference-plus-noise CPSD matrix by first estimating the interference PSD and afterwards the noise PSD in a maximum likelihood sense. Simulation experiments, where the proposed method is compared to other noise CPSD matrix estimators, show that it performs on par or better than competing methods, particularly, in situation where the interferenceto-noise ratio is large. Poul Hoang, Zheng-Hua Tan, Thomas Lunner, Jan Mark de Haan, Jesper Jensen 0001 |
ICASSP | 2 |
| 2020 | Adversarial Example Detection by Classification for Deep Speech RecognitionabstractMachine Learning systems are vulnerable to adversarial attacks and will highly likely produce incorrect outputs under these attacks. There are white-box and black-box attacks regarding to adversary's access level to the victim learning algorithm. To defend the learning systems from these attacks, existing methods in the speech domain focus on modifying input signals and testing the behaviours of speech recognizers. We, however, formulate the defense as a classification problem and present a strategy for systematically generating adversarial example datasets: one for white-box attacks and one for black-box attacks, containing both adversarial and normal examples. The white-box attack is a gradient-based method on Baidu DeepSpeech with the Mozilla Common Voice database while the black-box attack is a gradient-free method on a deep model-based keyword spotting system with the Google Speech Command dataset. The generated datasets are used to train a proposed Convolutional Neural Network (CNN), together with cepstral features, to detect adversarial examples. Experimental results show that, it is possible to accurately distinct between adversarial and normal examples for known attacks, in both single-condition and multi-condition training settings, while the performance degrades dramatically for unknown attacks. The adversarial datasets and the source code are made publicly available. Saeid Samizade, Zheng-Hua Tan, Chao Shen 0001, Xiaohong Guan |
ICASSP | 2 |
| 2020 | CC-Loss: Channel Correlation Loss for Image ClassificationabstractThe loss function is a key component in deep learning models. A commonly used loss function for classification is the cross entropy loss, which is a simple yet effective application of information theory for classification problems. Based on this loss, many other loss functions have been proposed, e.g., by adding intra-class and inter-class constraints to enhance the discriminative ability of the learned features. However, these loss functions fail to consider the connections between the feature distribution and the model structure. Aiming at addressing this problem, we propose a channel correlation loss (CC-Loss) that is able to constrain the specific relations between classes and channels as well as maintain the intra-class and the inter-class separability. CC-Loss uses a channel attention module to generate channel attention of features for each sample in the training stage. Next, an Euclidean distance matrix is calculated to make the channel attention vectors associated with the same class become identical and to increase the difference between different classes. Finally, we obtain a feature embedding with good intra-class compactness and inter-class separability. Experimental results show that two different backbone models trained with the proposed CC-Loss outperform the state-of-the-art loss functions on three image classification datasets. Dongliang Chang, Zhanyu Ma, Zheng-Hua Tan |
ICPR | 5 |
| 2020 | Vocoder-Based Speech Synthesis from Silent VideosabstractBoth acoustic and visual information influence human perception of speech. For this reason, the lack of audio in a video sequence determines an extremely low speech intelligibility for untrained lip readers. In this paper, we present a way to synthesise speech from the silent video of a talker using deep learning. The system learns a mapping function from raw video frames to acoustic features and reconstructs the speech with a vocoder synthesis algorithm. To improve speech reconstruction performance, our model is also trained to predict text information in a multi-task learning fashion and it is able to simultaneously reconstruct and recognise speech in real time. The results in terms of estimated speech quality and intelligibility show the effectiveness of our method, which exhibits an improvement over existing video-to-speech approaches. Daniel Michelsanti, Olga Slizovskaia, Gloria Haro, Emilia Gómez, Zheng-Hua Tan, Jesper Jensen 0001 |
INTERSPEECH | 5 |
| 2020 | rVAD: An unsupervised segment-based robust voice activity detection method
Zheng-Hua Tan, Achintya Kumar Sarkar, Najim Dehak |
Comput. Speech Lang. | 1 |
| 2020 | On Loss Functions for Supervised Monaural Time-Domain Speech EnhancementabstractMany deep learning-based speech enhancement algorithms are designed to minimize the mean-square error (MSE) in some transform domain between a predicted and a target speech signal. However, optimizing for MSE does not necessarily guarantee high speech quality or intelligibility, which is the ultimate goal of many speech enhancement algorithms. Additionally, only little is known about the impact of the loss function on the emerging class of time-domain deep learning-based speech enhancement systems. We study how popular loss functions influence the performance of time-domain deep learning-based speech enhancement systems. First, we demonstrate that perceptually inspired loss functions might be advantageous over classical loss functions like MSE. Furthermore, we show that the learning rate is a crucial design parameter even for adaptive gradient-based optimizers, which has been generally overlooked in the literature. Also, we found that waveform matching performance metrics must be used with caution as they in certain situations can fail completely. Finally, we show that a loss function based on scale-invariant signal-to-distortion ratio (SI-SDR) achieves good general performance across a range of popular speech enhancement evaluation metrics, which suggests that SI-SDR is a good candidate as a general-purpose loss function for speech enhancement systems. Morten Kolbæk, Zheng-Hua Tan, Søren Holdt Jensen, Jesper Jensen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Improved External Speaker-Robust Keyword Spotting for Hearing Assistive DevicesabstractFor certain applications, keyword spotting (KWS) requires some degree of personalization. This is the case for KWS for hearing assistive devices, e.g., hearing aids, where only the device user should be allowed to trigger the KWS system. In this paper, we first develop a new realistic hearing aid experimental framework. Next, using this framework we show that the performance of a state-of-the-art multi-task deep learning architecture exploiting cepstral features for joint KWS and users' own-voice/external speaker detection drops significantly. To overcome this problem, we use phase difference information through GCC-PHAT (Generalized Cross-Correlation with PHAse Transform)-based coefficients along with log-spectral magnitude features. In addition, we demonstrate that working in the perceptually-motivated constant-Q transform (CQT) domain instead of in the short-time Fourier transform (STFT) domain allows for the generation of compact and coherent features which provide superior KWS performance. Our experimental results show that our CQT-based proposal achieves a relative KWS accuracy improvement of around 18% compared to using cepstral features while dramatically decreasing the number of multiplications in the multi-task architecture, which is key in the context of low-resource devices like hearing assistive devices. Iván López-Espejo, Zheng-Hua Tan, Jesper Jensen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Online Multichannel Speech Enhancement Based on Recursive EM and DNN-Based Speech Presence EstimationabstractThis article presents a recursive expectation-maximization algorithm for online multichannel speech enhancement. A deep neural network mask estimator is used to compute the speech presence probability, which is then improved by means of statistical spatial models of the noisy speech and noise signals. The clean speech signal is estimated using beamforming, single-channel linear postfiltering and speech presence masking. The clean speech statistics and speech presence probabilities are finally used to compute the acoustic parameters for beamforming and postfiltering by means of maximum likelihood estimation. This iterative procedure is carried out on a frame-by-frame basis. The algorithm integrates the different estimates in a common statistical framework suitable for online scenarios. Moreover, our method can successfully exploit spectral, spatial and temporal speech properties. Our proposed algorithm is tested in different noisy environments using the multichannel recordings of the CHiME-4 database. The experimental results show that our method outperforms other related state-of-the-art approaches in noise reduction performance, while allowing low-latency processing for real-time applications. Juan M. Martín-Doñas, Jesper Jensen 0001, Zheng-Hua Tan, Ángel M. Gómez, Antonio M. Peinado |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | OSLNet: Deep Small-Sample Classification With an Orthogonal Softmax LayerabstractA deep neural network of multiple nonlinear layers forms a large function space, which can easily lead to overfitting when it encounters small-sample data. To mitigate overfitting in small-sample classification, learning more discriminative features from small-sample data is becoming a new trend. To this end, this paper aims to find a subspace of neural networks that can facilitate a large decision margin. Specifically, we propose the Orthogonal Softmax Layer (OSL), which makes the weight vectors in the classification layer remain orthogonal during both the training and test processes. The Rademacher complexity of a network using the OSL is only 1/K, where K is the number of classes, of that of a network using the fully connected classification layer, leading to a tighter generalization error bound. Experimental results demonstrate that the proposed OSL has better performance than the methods used for comparison on four small-sample benchmark datasets, as well as its applicability to large-sample datasets. Codes are available at: https://github.com/dongliangchang/OSLNet. Dongliang Chang, Zhanyu Ma, Zheng-Hua Tan, Jing-Hao Xue, Jie Cao 0014, Jingyi Yu 0001, Jun Guo 0002 |
IEEE Trans. Image Process. | 4 |
| 2020 | The Importance of Context When Recommending TV Content: Dataset and AlgorithmsabstractHome entertainment systems feature in a variety of usage scenarios with one or more simultaneous users, for whom the complexity of choosing media to consume has increased rapidly over the last decade. Users' decision processes are complex and highly influenced by contextual settings, but data supporting the development and evaluation of context-aware recommender systems are scarce. In this paper we present a dataset of self-reported TV consumption enriched with contextual information of viewing situations. We show how choice of genre associates with, among others, the number of present users and users' attention levels. Furthermore, we evaluate the performance of predicting chosen genres given different configurations of contextual information, and compare the results to contextless predictions. The results suggest that including contextual features in the prediction cause notable improvements, and both temporal and social context show significant contributions. Miklas S. Kristoffersen, Sven Ewan Shepstone, Zheng-Hua Tan |
IEEE Trans. Multim. | 3 |
| 2019 | Effects of Lombard Reflex on the Performance of Deep-learning-based Audio-visual Speech Enhancement SystemsabstractHumans tend to change their way of speaking when they are immersed in a noisy environment, a reflex known as Lombard effect. Current speech enhancement systems based on deep learning do not usually take into account this change in the speaking style, because they are trained with neutral (non-Lombard) speech utterances recorded under quiet conditions to which noise is artificially added. In this paper, we investigate the effects that the Lombard reflex has on the performance of audio-visual speech enhancement systems based on deep learning. The results show that a gap in the performance of as much as approximately 5 dB between the systems trained on neutral speech and the ones trained on Lombard speech exists. This indicates the benefit of taking into account the mismatch between neutral and Lombard speech in the design of audio-visual speech enhancement systems. Daniel Michelsanti, Zheng-Hua Tan, Sigurður Sigurðsson, Jesper Jensen 0001 |
ICASSP | 2 |
| 2019 | On Training Targets and Objective Functions for Deep-learning-based Audio-visual Speech EnhancementabstractAudio-visual speech enhancement (AV-SE) is the task of improving speech quality and intelligibility in a noisy environment using audio and visual information from a talker. Recently, deep learning techniques have been adopted to solve the AV-SE task in a supervised manner. In this context, the choice of the target, i.e. the quantity to be estimated, and the objective function, which quantifies the quality of this estimate, to be used for training is critical for the performance. This work is the first that presents an experimental study of a range of different targets and objective functions used to train a deep-learning-based AV-SE system. The results show that the approaches that directly estimate a mask perform the best overall in terms of estimated speech quality and intelligibility, although the model that directly estimates the log magnitude spectrum performs as good in terms of estimated speech quality. Daniel Michelsanti, Zheng-Hua Tan, Sigurður Sigurðsson, Jesper Jensen 0001 |
ICASSP | 2 |
| 2019 | Keyword Spotting for Hearing Assistive Devices Robust to External SpeakersabstractKeyword spotting (KWS) is experiencing an upswing due to the pervasiveness of small electronic devices that allow interaction with them via speech. Often, KWS systems are speaker-independent, which means that any person --user or not-- might trigger them. For applications like KWS for hearing assistive devices this is unacceptable, as only the user must be allowed to handle them. In this paper we propose KWS for hearing assistive devices that is robust to external speakers. A state-of-the-art deep residual network for small-footprint KWS is regarded as a basis to build upon. By following a multi-task learning scheme, this system is extended to jointly perform KWS and users' own-voice/external speaker detection with a negligible increase in the number of parameters. For experiments, we generate from the Google Speech Commands Dataset a speech corpus emulating hearing aids as a capturing device. Our results show that this multi-task deep residual network is able to achieve a KWS accuracy relative improvement of around 32% with respect to a system that does not deal with external speakers. Iván López-Espejo, Zheng-Hua Tan, Jesper Jensen 0001 |
INTERSPEECH | 2 |
| 2019 | Deep-learning-based audio-visual speech enhancement in presence of Lombard effect
Daniel Michelsanti, Zheng-Hua Tan, Sigurður Sigurðsson, Jesper Jensen 0001 |
Speech Commun. | 2 |
| 2019 | On the Relationship Between Short-Time Objective Intelligibility and Short-Time Spectral-Amplitude Mean-Square Error for Speech EnhancementabstractThe majority of deep neural network (DNN) based speech enhancement algorithms rely on the mean-square error (MSE) criterion of short-time spectral amplitudes (STSA), which has no apparent link to human perception, e.g., speech intelligibility. Short-time objective intelligibility (STOI), a popular state-of-the-art speech intelligibility estimator, on the other hand, relies on linear correlation of speech temporal envelopes. This raises the question if a DNN training criterion based on envelope linear correlation (ELC) can lead to improved speech intelligibility performance of DNN-based speech enhancement algorithms compared to algorithms based on the STSA-MSE criterion. In this paper, we derive that, under certain general conditions, the STSA-MSE and ELC criteria are practically equivalent, and we provide empirical data to support our theoretical results. Furthermore, our experimental findings suggest that the standard STSA minimum-MSE estimator is near optimal, if the objective is to enhance noisy speech in a manner, which is optimal with respect to the STOI speech intelligibility estimator. Morten Kolbæk, Zheng-Hua Tan, Jesper Jensen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Time-Contrastive Learning Based Deep Bottleneck Features for Text-Dependent Speaker VerificationabstractThere are a number of studies about extraction of bottleneck (BN) features from deep neural networks (DNNs) trained to discriminate speakers, pass-phrases, and triphone states for improving the performance of text-dependent speaker verification (TD-SV). However, a moderate success has been achieved. A recent study presented a time contrastive learning (TCL) concept to explore the non-stationarity of brain signals for classification of brain states. Speech signals have similar non-stationarity property, and TCL further has the advantage of having no need for labeled data. We therefore present a TCL based BN feature extraction method. The method uniformly partitions each speech utterance in a training dataset into a predefined number of multi-frame segments. Each segment in an utterance corresponds to one class, and class labels are shared across utterances. DNNs are then trained to discriminate all speech frames among the classes to exploit the temporal structure of speech. In addition, we propose a segment-based unsupervised clustering algorithm to re-assign class labels to the segments. TD-SV experiments were conducted on the RedDots challenge database. The TCL-DNNs were trained using speech data of fixed pass-phrases that were excluded from the TD-SV evaluation set, so the learned features can be considered phrase-independent. We compare the performance of the proposed TCL BN feature with those of short-time cepstral features and BN features extracted from DNNs discriminating speakers, pass-phrases, speaker+pass-phrase, as well as monophones whose labels and boundaries are generated by three different automatic speech recognition (ASR) systems. Experimental results show that the proposed TCL-BN outperforms cepstral features and speaker+pass-phrase discriminant BN features, and its performance is on par with those of ASR derived BN features. Moreover, the clustering method improves the TD-SV performance of TCL-BN and ASR derived BN features with respect to their standalone counterparts. We further study the TD-SV performance of fusing cepstral and BN features. Achintya Kumar Sarkar, Zheng-Hua Tan, Hao Tang 0002, Suwon Shon, James R. Glass |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | Monaural Speech Enhancement Using Deep Neural Networks by Maximizing a Short-Time Objective Intelligibility MeasureabstractIn this paper we propose a Deep Neural Network (D NN) based Speech Enhancement (SE) system that is designed to maximize an approximation of the Short-Time Objective Intelligibility (STOI) measure. We formalize an approximate-STOI cost function and derive analytical expressions for the gradients required for DNN training and show that these gradients have desirable properties when used together with gradient based optimization techniques. We show through simulation experiments that the proposed SE system achieves large improvements in estimated speech intelligibility, when tested on matched and unmatched natural noise types, at multiple signal-to-noise ratios. Furthermore, we show that the SE system, when trained using an approximate-STOI cost function performs on par with a system trained with a mean square error cost applied to short-time temporal envelopes. Finally, we show that the proposed SE system performs on par with a traditional DNN based Short- Time Spectral Amplitude (STSA) SE system in terms of estimated speech intelligibility. These results are important because they suggest that traditional DNN based STSA SE systems might be optimal in terms of estimated speech intelligibility. Morten Kolbæk, Zheng-Hua Tan, Jesper Jensen 0001 |
ICASSP | 2 |
| 2018 | Effectiveness of Single-Channel BLSTM Enhancement for Language IdentificationabstractThis paper proposes to apply deep neural network (DNN)-based single-channel speech enhancement (SE) to language identification. The 2017 language recognition evaluation (LRE17) introduced noisy audios from videos, in addition to the telephone conversation from past challenges. Because of that, adapting models from telephone speech to noisy speech from the video domain was required to obtain optimum performance. However, such adaptation requires knowledge of the audio domain and availability of in-domain data. Instead of adaptation, we propose to use a speech enhancement step to clean up the noisy audio as preprocessing for language identification. We used a bi-directional long short-term memory (BLSTM) neural network, which given log-Mel noisy features predicts a spectral mask indicating how clean each time-frequency bin is. The noisy spectrogram is multiplied by this predicted mask to obtain the enhanced magnitude spectrogram, and it is transformed back into the time domain by using the unaltered noisy speech phase. The experiments show significant improvement to language identification of noisy speech, for systems with and without domain adaptation, while preserving the identification performance in the telephone audio domain. In the best adapted state-of-the-art bottleneck i-vector system the relative improvement is 11.3% for noisy speech. Peter Sibbern Frederiksen, Jesús Villalba 0001, Shinji Watanabe 0001, Zheng-Hua Tan, Najim Dehak |
INTERSPEECH | 4 |
| 2018 | Public perception of android robots: Indications from an analysis of YouTube commentsabstractThe public perception of android robots is a field of growing applied relevance. Currently, most androids are confined within controlled environments rendering interactions between potential end-users, and robots challenging. Even more challenging is for researchers to investigate end-users' perception of androids. We exploit pre-existing YouTube comments as artifacts for quantitative content analysis to gain an indication of social perception on androids. We perform a content analysis of 10301 YouTube comments from four different videos, and reflect on the textual reactions to video stimuli of four extremely human-like android robots. We use text mining and machine learning techniques to process and analyze our corpus. Our findings reveal three equally important topics that should be considered for paving the way towards a robotic society: human-robot relationships, technical specifications, and the science fiction valley. Considering people's attitudes, fears and wishes towards androids, researchers can increase citizen awareness, and engagement. Evgenios Vlachos, Zheng-Hua Tan |
IROS | 2 |
| 2018 | The Sound or Silence: Investigating the Influence of Robot Noise on ProxemicsabstractIn the design of robots that have to share spaces with humans, design plays an important role in the acceptance, and sound is one further aspect that should not be neglected, in order to facilitate interaction and minimise repulsion. Following a previous research in which an uncomfortable noise from a robot was identified, we present an interaction study in which the same noise was tested in order to measure the effect on proxemics, as opposed to a silent condition and to a different version of the sound, masked with ambient music. The results obtained with the participation of 60 subjects show the effectiveness of the mask in avoiding the negative effect of the noise. Gabriele Trovato, Renato Paredes, Javier Balvin, Francisco Cuéllar, Nicolai Bæk Thomsen, Soren Bech, Zheng-Hua Tan |
RO-MAN | 7 |
| 2018 | A Dataset for Inferring Contextual Preferences of Users Watching TVabstractStudies have shown that contextual settings play an important role in users' decision processes of what to consume, but data supporting the investigation of context-aware recommender systems are scarce. In this paper we present a TV consumption dataset enriched with contextual information of viewing situations. The dataset is designed for studying the intrinsic complexity of TV watching activities, and hence we also evaluate the performance of predicting chosen genres given contextual settings, and compare the results to contextless predictions. The results suggest a significant improvement by including contextual features in the prediction. Miklas S. Kristoffersen, Sven Ewan Shepstone, Zheng-Hua Tan |
UMAP | 3 |
| 2018 | Incorporating pass-phrase dependent background models for text-dependent speaker verification
Achintya Kumar Sarkar, Zheng-Hua Tan |
Comput. Speech Lang. | 2 |
| 2018 | Latent Dirichlet mixture model
Jen-Tzung Chien, Chao-Hsi Lee, Zheng-Hua Tan |
Neurocomputing | 3 |
| 2018 | Recent advances in machine learning for non-Gaussian data processing
Zhanyu Ma, Jen-Tzung Chien, Zheng-Hua Tan, Yi-Zhe Song, Jalil Taghia, Ming Xiao 0001 |
Neurocomputing | 3 |
| 2018 | A spatial self-similarity based feature learning method for face recognition under varying poses
Xiaodong Duan, Zheng-Hua Tan |
Pattern Recognit. Lett. | 2 |
| 2018 | Refinement and validation of the binaural short time objective intelligibility measure for spatially diverse conditions
Asger Heidemann Andersen, Jan Mark de Haan, Zheng-Hua Tan, Jesper Jensen 0001 |
Speech Commun. | 3 |
| 2018 | A perceptually motivated LP residual estimator in noisy and reverberant environments
Renhua Peng, Zheng-Hua Tan, Xiaodong Li 0002, Chengshi Zheng |
Speech Commun. | 2 |
| 2018 | Audio-Based Granularity-Adapted Emotion ClassificationabstractThis paper introduces a novel framework for combining the strengths of machine-based and human-based emotion classification. Peoples' ability to tell similar emotions apart is known as emotional granularity, which can be high or low, and is measurable. This paper proposes granularity-adapted classification that can be used as a front-end to drive a recommender, based on emotions from speech. In this context, incorrectly predicted peoples' emotions could lead to poor recommendations, reducing user satisfaction. Instead of identifying a single emotion class, an adapted class is proposed, and is an aggregate of underlying emotion classes chosen based on granularity. In the recommendation context, the adapted class maps to a larger region in valence-arousal space, from which a list of potentially more similar content items is drawn, and recommended to the user. To determine the effectiveness of adapted classes, we measured the emotional granularity of subjects, and for each subject, used their pairwise similarity judgments of emotion to compare the effectiveness of adapted classes versus single emotion classes taken from a baseline system. A customized Euclidean-based similarity metric is used to measure the relative proximity of emotion classes. Results show that granularity-adapted classification can improve the potential similarity by up to 9.6 percent. Sven Ewan Shepstone, Zheng-Hua Tan, Søren Holdt Jensen |
IEEE Trans. Affect. Comput. | 2 |
| 2018 | Nonintrusive Speech Intelligibility Prediction Using Convolutional Neural NetworksabstractSpeech Intelligibility Prediction (SIP) algorithms are becoming popular tools within the development and operation of speech processing devices and algorithms. However, many SIP algorithms require knowledge of the underlying clean speech; a signal that is often not available in real-world applications. This has led to increased interest in nonintrusive SIP algorithms, which do not require clean speech to make predictions. In this paper, we investigate the use of Convolutional Neural Networks (CNNs) for nonintrusive SIP. To do so, we utilize a CNN architecture that shows similarities to existing SIP algorithms, in terms of computational structure, and which allows for easy and meaningful visualization and interpretation of trained weights. We evaluate this architecture using a large dataset obtained by combining datasets from the literature. The proposed method shows high prediction performance when compared with four existing intrusive and nonintrusive SIP algorithms. This demonstrates the potential of deep learning for speech intelligibility prediction. Asger Heidemann Andersen, Jan Mark de Haan, Zheng-Hua Tan, Jesper Jensen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Bias-Compensated Informed Sound Source Localization Using Relative Transfer FunctionsabstractIn this paper, we consider the problem of estimating the target sound direction of arrival (DoA) for a hearing aid (HA) system, which can connect to a wireless microphone worn by the talker of interest. The wireless microphone “informs” the HA system about the noise-free target speech. To estimate the DoA, we consider a maximum-likelihood approach, and we assume that a database of DoA-dependent relative transfer functions (RTFs) has been measured in advance and is available. The proposed DoA estimator is able to take the available noise-free target speech, ambient noise characteristics, and the shadowing effect of the user's head on the received signals into account, and it supports both monaural and binaural microphone array configurations. Moreover, we analytically analyze the bias in the proposed estimator and introduce a modified estimator, which has been compensated for the bias. We demonstrate that the proposed method has lower computational complexity and better performance than recent RTF-based estimators. Furthermore, to decrease the number of parameters required to be wirelessly exchanged between the HAs in binaural configurations, we propose an information fusion strategy, which avoids transmitting microphone signals between the HAs. An important benefit of the proposed IF strategy is that the number of parameters to be exchanged between the HAs is independent of the number of HA microphones. Finally, we investigate the performance of variants of the proposed estimator extensively in different noisy and reverberant situations. Mojtaba Farmani, Michael Syskind Pedersen, Zheng-Hua Tan, Jesper Jensen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Robust Voice Liveness Detection and Speaker Verification Using Throat MicrophonesabstractWhile having a wide range of applications, automatic speaker verification (ASV) systems are vulnerable to spoofing attacks, in particular, replay attacks that are effective and easy to implement. Most prior work on detecting replay attacks uses audio from a single acoustic microphone only, leading to difficulties in detecting high-end replay attacks close to indistinguishable from live human speech. In this paper, we study the use of a special body-conducted sensor, throat microphone (TM), for combined voice liveness detection (VLD) and ASV in order to improve both robustness and security of ASV against replay attacks. We first investigate the possibility and methods of attacking a TM-based ASV system, followed by a pilot data collection. Second, we study the use of spectral features for VLD using both single-channel and dual-channel ASV systems. We carry out speaker verification experiments using Gaussian mixture model with universal background model (GMM-UBM) and i-vector based systems on a dataset of 38 speakers collected by us. We have achieved considerable improvement in recognition accuracy, with the use of dual-microphone setup. In experiments with noisy test speech, the false acceptance rate (FAR) of the dual-microphone GMM-UBM based system for recorded speech reduces from 69.69% to 18.75%. The FAR of replay condition further drops to 0% when this dual-channel ASV system is integrated with the new dual-channel voice liveness detector. Md. Sahidullah, Dennis Alexander Lehmann Thomsen, Rosa González Hautamäki, Tomi Kinnunen, Zheng-Hua Tan, Robert Parts, Martti Pitkänen |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2018 | Decorrelation of Neutral Vector Variables: Theory and ApplicationsabstractIn this paper, we propose novel strategies for neutral vector variable decorrelation. Two fundamental invertible transformations, namely, serial nonlinear transformation and parallel nonlinear transformation, are proposed to carry out the decorrelation. For a neutral vector variable, which is not multivariate-Gaussian distributed, the conventional principal component analysis cannot yield mutually independent scalar variables. With the two proposed transformations, a highly negatively correlated neutral vector can be transformed to a set of mutually independent scalar variables with the same degrees of freedom. We also evaluate the decorrelation performances for the vectors generated from a single Dirichlet distribution and a mixture of Dirichlet distributions. The mutual independence is verified with the distance correlation measurement. The advantages of the proposed decorrelation strategies are intensively studied and demonstrated with synthesized data and practical application evaluations. Zhanyu Ma, Jing-Hao Xue, Arne Leijon, Zheng-Hua Tan, Zhen Yang 0004, Jun Guo 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2018 | Spoofing Detection in Automatic Speaker Verification Systems Using DNN Classifiers and Dynamic Acoustic FeaturesabstractWith the development of speech synthesis technology, automatic speaker verification (ASV) systems have encountered the serious challenge of spoofing attacks. In order to improve the security of ASV systems, many antispoofing countermeasures have been developed. In the front-end domain, much research has been conducted on finding effective features which can distinguish spoofed speech from genuine speech and the published results show that dynamic acoustic features work more effectively than static ones. In the back-end domain, Gaussian mixture model (GMM) and deep neural networks (DNNs) are the two most popular types of classifiers used for spoofing detection. The log-likelihood ratios (LLRs) generated by the difference of human and spoofing log-likelihoods are used as spoofing detection scores. In this paper, we train a five-layer DNN spoofing detection classifier using dynamic acoustic features and propose a novel, simple scoring method only using human log-likelihoods (HLLs) for spoofing detection. We mathematically prove that the new HLL scoring method is more suitable for the spoofing detection task than the classical LLR scoring method, especially when the spoofing speech is very similar to the human speech. We extensively investigate the performance of five different dynamic filter bank-based cepstral features and constant Q cepstral coefficients (CQCC) in conjunction with the DNN-HLL method. The experimental results show that, compared to the GMM-LLR method, the DNN-HLL method is able to significantly improve the spoofing detection accuracy. Compared with the CQCC-based GMM-LLR baseline, the proposed DNN-HLL model reduces the average equal error rate of all attack types to 0.045%, thus exceeding the performance of previously published approaches for the ASVspoof 2015 Challenge task. Fusing the CQCC-based DNN-HLL spoofing detection system with ASV systems, the false acceptance rate on spoofing attacks can be reduced significantly. Hong Yu 0006, Zheng-Hua Tan, Zhanyu Ma, Rainer Martin 0001, Jun Guo 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | A non-intrusive Short-Time Objective Intelligibility measureabstractWe propose a non-intrusive intelligibility measure for noisy and non-linearly processed speech, i.e. a measure which can predict intelligibility from a degraded speech signal without requiring a clean reference signal. The proposed measure is based on the Short-Time Objective Intelligibility (STOI) measure. In particular, the non-intrusive STOI measure estimates clean signal amplitude envelopes from the degraded signal. Subsequently, the STOI measure is evaluated by use of the envelopes of the degraded signal and the estimated clean envelopes. The performance of the proposed measure is evaluated on a dataset including speech in different noise types, processed with binary masks. The measure is shown to predict intelligibility well in all tested conditions, with the exception of those including a single competing speaker. While the measure does not perform as well as the original (intrusive) STOI measure, it is shown to outperform existing non-intrusive measures. Asger Heidemann Andersen, Jan Mark de Haan, Zheng-Hua Tan, Jesper Jensen 0001 |
ICASSP | 3 |
| 2017 | RedDots replayed: A new replay spoofing attack corpus for text-dependent speaker verification researchabstractThis paper describes a new database for the assessment of automatic speaker verification (ASV) vulnerabilities to spoofing attacks. In contrast to other recent data collection efforts, the new database has been designed to support the development of replay spoofing countermeasures tailored towards the protection of text-dependent ASV systems from replay attacks in the face of variable recording and playback conditions. Derived from the re-recording of the original RedDots database, the effort is aligned with that in text-dependent ASV and thus well positioned for future assessments of replay spoofing countermeasures, not just in isolation, but in integration with ASV. The paper describes the database design and re-recording, a protocol and some early spoofing detection results. The new “RedDots Replayed” database is publicly available through a creative commons license. Tomi Kinnunen, Md. Sahidullah, Mauro Falcone, Luca Costantini, Rosa González Hautamäki, Dennis Alexander Lehmann Thomsen, Achintya Kumar Sarkar, Zheng-Hua Tan, Héctor Delgado, Massimiliano Todisco, Nicholas W. D. Evans, Ville Hautamäki, Kong-Aik Lee |
ICASSP | 8 |
| 2017 | Permutation invariant training of deep models for speaker-independent multi-talker speech separationabstractWe propose a novel deep learning training criterion, named permutation invariant training (PIT), for speaker independent multi-talker speech separation, commonly known as the cocktail-party problem. Different from the multi-class regression technique and the deep clustering (DPCL) technique, our novel approach minimizes the separation error directly. This strategy effectively solves the long-lasting label permutation problem, that has prevented progress on deep learning based techniques for speech separation. We evaluated PIT on the WSJ0 and Danish mixed-speech separation tasks and found that it compares favorably to non-negative matrix factorization (NMF), computational auditory scene analysis (CASA), and DPCL and generalizes well over unseen speakers and languages. Since PIT is simple to implement and can be easily integrated and combined with other advanced techniques, we believe improvements built upon PIT can eventually solve the cocktail-party problem. Dong Yu 0001, Morten Kolbæk, Zheng-Hua Tan, Jesper Jensen 0001 |
ICASSP | 3 |
| 2017 | Weighted Score Based Fast Converging CO-training with Application to Audio-Visual Person IdentificationabstractOne potential problem in real classification applications is that the amount of labeled training data is insufficient since it is usually time-consuming to label data manually. When multiple modalities are available, it is possible to train an initial classifier for each modality using a small amount of labeled data, and then re-train each classifier using unlabeled data associated with the labels generated from the other modalities. This can be achieved by the well-known CO-training algorithm. Assuming that two modalities are available, it only takes the information from the other modality but not that from the self modality into account when choosing data, which usually results in slow convergence of classification accuracy. This may make the CO-training procedure time-consuming. To overcome this, we present a novel modification to the original CO-training algorithm, which is concerned with how new samples are chosen at each iteration to re-train the classifiers in order to improve the convergence of classification accuracy. In our method, the new data is chosen based on the weighted scores which are generated from both modalities instead of only the scores from the other modality as in the original CO-training. We apply both the modified and original CO-training methods on multi-modal person identification task using speech and vision. Experiments on a publicly available database show that our method outperforms the original CO-training by a large margin, in terms of convergence of classification accuracy on a separate testing data set. Xiaodong Duan, Nicolai Bæk Thomsen, Zheng-Hua Tan, Børge Lindberg, Søren Holdt Jensen |
ICTAI | 3 |
| 2017 | On the Use of Band Importance Weighting in the Short-Time Objective Intelligibility MeasureabstractSpeech intelligibility prediction methods are popular tools within the speech processing community for objective evaluation of speech intelligibility of e.g.enhanced speech.The Short-Time Objective Intelligibility (STOI) measure has become highly used due to its simplicity and high prediction accuracy.In this paper we investigate the use of Band Importance Functions (BIFs) in the STOI measure, i.e. of unequally weighting the contribution of speech information from each frequency band.We do so by fitting BIFs to several datasets of measured intelligibility, and cross evaluating the prediction performance.Our findings indicate that it is possible to improve prediction performance in specific situations.However, it has not been possible to find BIFs which systematically improve prediction performance beyond the data used for fitting.In other words, we find no evidence that the performance of the STOI measure can be improved considerably by extending it with a non-uniform BIF. Asger Heidemann Andersen, Jan Mark de Haan, Zheng-Hua Tan, Jesper Jensen 0001 |
INTERSPEECH | 3 |
| 2017 | The I4U Mega Fusion and Collaboration for NIST Speaker Recognition Evaluation 2016abstract18th Annual Conference of the International Speech Communication Association, INTERSPEECH 2017, Stockholm, Sweden, 20-24 August 2017 Kong-Aik Lee, Ville Hautamäki, Tomi Kinnunen, Anthony Larcher, Andreas Nautsch, Themos Stafylakis, Gang Liu 0001, Mickael Rouvier, Wei Rao 0002, Federico Alegre, Man-Wai Mak, Achintya Kumar Sarkar, Héctor Delgado, Rahim Saeidi, Hagai Aronowitz, Aleksandr Sizov, Hanwu Sun, Trung Hieu Nguyen 0001, Guangsen Wang, Bin Ma 0001, Ville Vestman, Md. Sahidullah, M. Halonen, Anssi Kanervisto, Gaël Le Lan, Fahimeh Bahmaninezhad, Sergey Isadskiy, Christian Rathgeb, Christoph Busch 0001, Georgios Tzimiropoulos, Q. Qian, Q. Zhao, J. Xue, R. Jin, T. Zhao, Pierre-Michel Bousquet, Moez Ajili, Waad Ben Kheder, Driss Matrouf, Zhi Hao Lim, Chenglin Xu, Haihua Xu 0001, Chng Eng Siong, Benoit G. B. Fauve, Kaavya Sriskandaraja, Vidhyasaharan Sethu, W. W. Lin, Dennis Alexander Lehmann Thomsen, Zheng-Hua Tan, Massimiliano Todisco, Nicholas W. D. Evans, Haizhou Li 0001, John H. L. Hansen, Jean-François Bonastre, Eliathamby Ambikairajah |
INTERSPEECH | 56 |
| 2017 | Conditional Generative Adversarial Networks for Speech Enhancement and Noise-Robust Speaker VerificationabstractImproving speech system performance in noisy environments remains a challenging task, and speech enhancement (SE) is one of the effective techniques to solve the problem.Motivated by the promising results of generative adversarial networks (GANs) in a variety of image processing tasks, we explore the potential of conditional GANs (cGANs) for SE, and in particular, we make use of the image processing framework proposed by Isola et al. [1] to learn a mapping from the spectrogram of noisy speech to an enhanced counterpart.The SE cGAN consists of two networks, trained in an adversarial manner: a generator that tries to enhance the input noisy spectrogram, and a discriminator that tries to distinguish between enhanced spectrograms provided by the generator and clean ones from the database using the noisy spectrogram as a condition.We evaluate the performance of the cGAN method in terms of perceptual evaluation of speech quality (PESQ), short-time objective intelligibility (STOI), and equal error rate (EER) of speaker verification (an example application).Experimental results show that the cGAN method overall outperforms the classical short-time spectral amplitude minimum mean square error (STSA-MMSE) SE algorithm, and is comparable to a deep neural network-based SE approach (DNN-SE). Daniel Michelsanti, Zheng-Hua Tan |
INTERSPEECH | 2 |
| 2017 | Improving Speaker Verification Performance in Presence of Spoofing Attacks Using Out-of-Domain Spoofed DataabstractAutomatic speaker verification (ASV) systems are vulnerable to spoofing attacks using speech generated by voice conversion and speech synthesis techniques. Commonly, a countermeasure (CM) system is integrated with an ASV system for improved protection against spoofing attacks. But integration of the two systems is challenging and often leads to increased false rejection rates. Furthermore, the performance of CM severely degrades if in-domain development data are unavailable. In this study, therefore, we propose a solution that uses two separate background models – one from human speech and another from spoofed data. During test, the ASV score for an input utterance is computed as the difference of the log-likelihood against the target model and the combination of the log-likelihoods against two background models. Evaluation experiments are conducted using the joint ASV and CM protocol of ASVspoof 2015 corpus consisting of text-independent ASV tasks with short utterances. Our proposed system reduces error rates in the presence of spoofing attacks by using out-of-domain spoofed data for system development, while maintaining the performance for zero-effort imposter attacks compared to the baseline system. Achintya Kumar Sarkar, Md. Sahidullah, Zheng-Hua Tan, Tomi Kinnunen |
INTERSPEECH | 3 |
| 2017 | Adversarial Network Bottleneck Features for Noise Robust Speaker VerificationabstractIn this paper, we propose a noise robust bottleneck feature representation which is generated by an adversarial network (AN).The AN includes two cascade connected networks, an encoding network (EN) and a discriminative network (DN).Melfrequency cepstral coefficients (MFCCs) of clean and noisy speech are used as input to the EN and the output of the EN is used as the noise robust feature.The EN and DN are trained in turn, namely, when training the DN, noise types are selected as the training labels and when training the EN, all labels are set as the same, i.e., the clean speech label, which aims to make the AN features invariant to noise and thus achieve noise robustness.We evaluate the performance of the proposed feature on a Gaussian Mixture Model-Universal Background Model based speaker verification system, and make comparison to MFCC features of speech enhanced by short-time spectral amplitude minimum mean square error (STSA-MMSE) and deep neural network-based speech enhancement (DNN-SE) methods.Experimental results on the RSR2015 database show that the proposed AN bottleneck feature (AN-BN) dramatically outperforms the STSA-MMSE and DNN-SE based MFCCs for different noise types and signal-to-noise ratios.Furthermore, the AN-BN feature is able to improve the speaker verification performance under the clean condition. Hong Yu 0006, Zheng-Hua Tan, Zhanyu Ma, Jun Guo 0002 |
INTERSPEECH | 2 |
| 2017 | Informed Sound Source Localization Using Relative Transfer Functions for Hearing Aid ApplicationsabstractRecent hearing aid systems (HASs) can connect to a wireless microphone worn by the talker of interest. This feature gives the HASs access to a noise-free version of the target signal. In this paper, we address the problem of estimating the target sound direction of arrival (DoA) for a binaural HAS given access to the noise-free content of the target signal. To estimate the DoA, we present a maximum-likelihood framework which takes the shadowing effect of the user's head on the received signals into account by modeling the relative transfer functions (RTFs) between the HAS's microphones. We propose three different RTF models which have different degrees of accuracy and individualization. Furthermore, we show that the proposed DoA estimators can be formulated in terms of inverse discrete Fourier transforms to evaluate the likelihood function computationally efficiently. We extensively assess the performance of the proposed DoA estimators for various DoAs, signal to noise ratios, and in different noisy and reverberant situations. The results show that the proposed estimators improve the performance markedly over other recently proposed “informed” DoA estimator. Mojtaba Farmani, Michael Syskind Pedersen, Zheng-Hua Tan, Jesper Jensen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | Speech Intelligibility Potential of General and Specialized Deep Neural Network Based Speech Enhancement SystemsabstractIn this paper, we study aspects of single microphone speech enhancement (SE) based on deep neural networks (DNNs). Specifically, we explore the generalizability capabilities of state-of-the-art DNN-based SE systems with respect to the background noise type, the gender of the target speaker, and the signal-to-noise ratio (SNR). Furthermore, we investigate how specialized DNN-based SE systems, which have been trained to be either noise type specific, speaker specific or SNR specific, perform relative to DNN based SE systems that have been trained to be noise type general, speaker general, and SNR general. Finally, we compare how a DNN-based SE system trained to be noise type general, speaker general, and SNR general performs relative to a state-of-the-art short-time spectral amplitude minimum mean square error (STSA-MMSE) based SE algorithm. We show that DNN-based SE systems, when trained specifically to handle certain speakers, noise types and SNRs, are capable of achieving large improvements in estimated speech quality (SQ) and speech intelligibility (SI), when tested in matched conditions. Furthermore, we show that improvements in estimated SQ and SI can be achieved by a DNN-based SE system when exposed to unseen speakers, genders and noise types, given a large number of speakers and noise types have been used in the training of the system. In addition, we show that a DNN-based SE system that has been trained using a large number of speakers and a wide range of noise types outperforms a state-of-the-art STSA-MMSE based SE method, when tested using a range of unseen speakers and noise types. Finally, a listening test using several DNN-based SE systems tested in unseen speaker conditions show that these systems can improve SI for some SNR and noise type configurations but degrade SI for others. Morten Kolbæk, Zheng-Hua Tan, Jesper Jensen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Multitalker Speech Separation With Utterance-Level Permutation Invariant Training of Deep Recurrent Neural NetworksabstractIn this paper, we propose the utterance-level permutation invariant training (uPIT) technique. uPIT is a practically applicable, end-to-end, deep-learning-based solution for speaker independent multitalker speech separation. Specifically, uPIT extends the recently proposed permutation invariant training (PIT) technique with an utterance-level cost function, hence eliminating the need for solving an additional permutation problem during inference, which is otherwise required by frame-level PIT. We achieve this using recurrent neural networks (RNNs) that, during training, minimize the utterance-level separation error, hence forcing separated frames belonging to the same speaker to be aligned to the same output stream. In practice, this allows RNNs, trained with uPIT, to separate multitalker mixed speech without any prior knowledge of signal duration, number of speakers, speaker identity, or gender. We evaluated uPIT on the WSJ0 and Danish two- and three-talker mixed-speech separation tasks and found that uPIT outperforms techniques based on nonnegative matrix factorization and computational auditory scene analysis, and compares favorably with deep clustering, and the deep attractor network. Furthermore, we found that models trained with uPIT generalize well to unseen speakers and languages. Finally, we found that a single model, trained with uPIT, can handle both two-speaker, and three-speaker speech mixtures. Morten Kolbæk, Dong Yu 0001, Zheng-Hua Tan, Jesper Jensen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Concurrent localization of sound sources and dual-microphone sub-arrays using TOFs
Mojtaba Farmani, Richard Heusdens, Michael Syskind Pedersen, Zheng-Hua Tan, Jesper Jensen 0001 |
FUSION | 4 |
| 2016 | A method for predicting the intelligibility of noisy and non-linearly enhanced binaural speechabstractWe propose and evaluate a binaural speech intelligibility measure. The measure is a binaural extension of the Short-Time Objective Intelligibility (STOI) measure and focuses on predicting the intelligibility of noisy speech which has been enhanced by a speech processing algorithm (e.g. in a hearing aid). We show that the measure can accurately predict 1) the Speech Reception Threshold (SRT) for a frontal speaker masked by a point noise source in the horizontal plane, 2) the improvement in SRT obtained by independently processing the left and right ear signals with Ideal Time Frequency Segregation (ITFS), and 3) the intelligibility of speech in the presence of multiple interferers as well as the effect of processing the noisy signals with 2-microphone MVDR beamforming as used in hearing aids. Finally, we show that the computational demands associated with the measure are favourable in comparison with those of a previously proposed measure with similar properties. Asger Heidemann Andersen, Jan Mark de Haan, Zheng-Hua Tan, Jesper Jensen 0001 |
ICASSP | 3 |
| 2016 | Informed Direction of Arrival estimation using a spherical-head model for Hearing Aid applicationsabstractIn this paper, we propose a Direction of Arrival (DoA) estimator for a Hearing Aid System (HAS) which can connect to a wireless microphone worn by a target talker. The wireless microphone "informs" the HAS about the almost noise-free content of the target sound, and the proposed DoA estimator uses the knowledge of the noise-free target sound and the received microphone signals to estimate the DoA via a maximum likelihood approach. Moreover, the proposed DoA estimator resorts to a user-independent spherical-head model to consider the acoustic impacts of the head on the received signals at the HAS. Further, the proposed DoA estimator uses an Inverse Discrete Fourier Transform (IDFT) technique to evaluate the likelihood function computationally efficiently. We assessed the performance of the proposed estimator for various DoAs, Signal to Noise Ratios (SNRs), and target distances in different noisy and reverberant situations. The proposed estimator improves the performance markedly over other recently proposed "informed" DoA estimators. Mojtaba Farmani, Michael Syskind Pedersen, Zheng-Hua Tan, Jesper Jensen 0001 |
ICASSP | 3 |
| 2016 | Adaptive overcurrent protection for microgrids in extensive distribution systemsabstractMicrogrid is regarded as a new form to integrate the increasing penetration of distributed generation units (DGs) in the extensive distribution systems. This paper proposes an adaptive overcurrent protection strategy for a microgrid network. The protection coordination of the overcurrent relays is treated as a linear programming problem for the different operation states. In the control center, an artificial neural network (ANN) model is trained with real-time measurements to identify the states whether there is a fault on the line segment. Fault location is estimated further with the same measurements in another neural network model. Reconfigurations can be performed to modify the settings of the on-field relays to enhance the reliable operation for the different operational situations. The test results show that the adaptive overcurrent protection scheme with the assistance of estimation model can modify the protective settings for the new operation state accurately and intelligently. Hengwei Lin, Josep M. Guerrero, Chenxi Jia, Zheng-Hua Tan, Juan C. Vasquez 0001, Chengxi Liu |
IECON | 4 |
| 2016 | Utterance Verification for Text-Dependent Speaker Recognition: A Comparative Assessment Using the RedDots CorpusabstractText-dependent automatic speaker verification naturally calls for the simultaneous verification of speaker identity and spoken content. These two tasks can be achieved with automatic speaker verification (ASV) and utterance verification (UV) technologies. While both have been addressed previously in the literature, a treatment of simultaneous speaker and utterance verification with a modern, standard database is so far lacking. This is despite the burgeoning demand for voice biometrics in a plethora of practical security applications. With the goal of improving overall verification performance, this paper reports different strategies for simultaneous ASV and UV in the context of short-duration, text-dependent speaker verification. Experiments performed on the recently released RedDots corpus are reported for three different ASV systems and four different UV systems. Results show that the combination of utterance verification with automatic speaker verification is (almost) universally beneficial with significant performance improvements being observed. Tomi Kinnunen, Md. Sahidullah, Ivan Kukanov, Héctor Delgado, Massimiliano Todisco, Achintya Kumar Sarkar, Nicolai Bæk Thomsen, Ville Hautamäki, Nicholas W. D. Evans, Zheng-Hua Tan |
INTERSPEECH | 10 |
| 2016 | HAPPY Team Entry to NIST OpenSAD Challenge: A Fusion of Short-Term Unsupervised and Segment i-Vector Based Speech Activity DetectorsabstractSpeech activity detection (SAD), the task of locating speech segments from a given recording, remains challenging under acoustically degraded conditions. In 2015, National Institute of Standards and Technology (NIST) coordinated OpenSAD bench-mark. We summarize “HAPPY” team effort to OpenSAD. SADs come in both unsupervised and supervised flavors, the latter requiring a labeled training set. Our solution fuses six base SADs (2 supervised and 4 unsupervised). The individually best SAD, in terms of detection cost function (DCF), is supervised and uses adaptive segmentation with i-vectors to represent the segments. Fusion of the six base SADs yields a relative decrease of 9.3% in DCF over this SAD. Further, relative decrease of 17.4% is obtained by incorporating channel detection side information. Tomi Kinnunen, Alexey Sholokhov, Elie Khoury 0001, Dennis Alexander Lehmann Thomsen, Md. Sahidullah, Zheng-Hua Tan |
INTERSPEECH | 6 |
| 2016 | Integrated Spoofing Countermeasures and Automatic Speaker Verification: An Evaluation on ASVspoof 2015abstractIt is well known that automatic speaker verification (ASV) systems can be vulnerable to spoofing. The community has responded to the threat by developing dedicated countermeasures aimed at detecting spoofing attacks. Progress in this area has accelerated over recent years, partly as a result of the first standard evaluation, ASVspoof 2015, which focused on spoofing detection in isolation from ASV. This paper investigates the integration of state-of-the-art spoofing countermeasures in combination with ASV. Two general strategies to countermeasure integration are reported: cascaded and parallel. The paper reports the first comparative evaluation of each approach performed with the ASVspoof 2015 corpus. Results indicate that, even in the case of varying spoofing attack algorithms, ASV performance remains robust when protected with a diverse set of integrated countermeasures. Md. Sahidullah, Héctor Delgado, Massimiliano Todisco, Hong Yu 0015, Tomi Kinnunen, Nicholas W. D. Evans, Zheng-Hua Tan |
INTERSPEECH | 7 |
| 2016 | Robust Speaker Recognition with Combined Use of Acoustic and Throat Microphone SpeechabstractAccuracy of automatic speaker recognition (ASV) systems degrades severely in the presence of background noise. In this paper, we study the use of additional side information provided by a body-conducted sensor, throat microphone. Throat microphone signal is much less affected by background noise in comparison to acoustic microphone signal. This makes throat microphones potentially useful for feature extraction or speech activity detection. This paper, firstly, proposes a new prototype system for simultaneous data-acquisition of acoustic and throat microphone signals. Secondly, we study the use of this additional information for both speech activity detection, feature extraction and fusion of the acoustic and throat microphone signals. We collect a pilot database consisting of 38 subjects including both clean and noisy sessions. We carry out speaker verification experiments using Gaussian mixture model with universal background model (GMM-UBM) and i-vector based system. We have achieved considerable improvement in recognition accuracy even in highly degraded conditions. Md. Sahidullah, Rosa González Hautamäki, Dennis Alexander Lehmann Thomsen, Tomi Kinnunen, Zheng-Hua Tan, Ville Hautamäki, Robert Parts, Martti Pitkänen |
INTERSPEECH | 5 |
| 2016 | Text Dependent Speaker Verification Using Un-Supervised HMM-UBM and Temporal GMM-UBMabstractIn this paper, we investigate the Hidden Markov Model (HMM) and the temporal Gaussian Mixture Model (GMM) systems based on the Universal Background Model (UBM) concept to capture temporal information of speech for Text Dependent (TD) Speaker Verification (SV). In TD-SV, target speakers are constrained to use only predefined fixed sentence/s during both the enrollment and the test process. The temporal information is therefore important in the sense of utterance verification, i.e. whether the test utterance has the same sequence of textual content as the utterance used during the target enrollment. However, the temporal information is not considered in the classical GMM-UBM based TD-SV system. Moreover, no transcription knowledge of the speech is required in the HMM-UBM and temporal GMM-UBM based systems. We also study the fusion of the HMM-UBM, the temporal GMM-UBM and the classical GMM-UBM systems in SV. We show that the HMM-UBM system yields better performance than the other systems in most cases. Further, fusion of the systems improve the overall speaker verification performance. The results are shown in the different tasks of the RedDots challenge 2016 database. Achintya Kumar Sarkar, Zheng-Hua Tan |
INTERSPEECH | 2 |
| 2016 | Speaker-Dependent Dictionary-Based Speech Enhancement for Text-Dependent Speaker VerificationabstractThe problem of text-dependent speaker verification under noisy conditions is becoming ever more relevant, due to increased usage for authentication in real-world applications. Classical methods for noise reduction such as spectral subtraction and Wiener filtering introduce distortion and do not perform well in this setting. In this work we compare the performance of different noise reduction methods under different noise conditions in terms of speaker verification when the text is known and the system is trained on clean data (mis-matched conditions). We furthermore propose a new approach based on dictionary-based noise reduction and compare it to the baseline methods. Nicolai Bæk Thomsen, Dennis Alexander Lehmann Thomsen, Zheng-Hua Tan, Børge Lindberg, Søren Holdt Jensen |
INTERSPEECH | 3 |
| 2016 | Further optimisations of constant Q cepstral processing for integrated utterance and text-dependent speaker verificationabstractMany authentication applications involving automatic speaker verification (ASV) demand robust performance using short-duration, fixed or prompted text utterances. Text constraints not only reduce the phone-mismatch between enrolment and test utterances, which generally leads to improved performance, but also provide an ancillary level of security. This can take the form of explicit utterance verification (UV). An integrated UV + ASV system should then verify access attempts which contain not just the expected speaker, but also the expected text content. This paper presents such a system and introduces new features which are used for both UV and ASV tasks. Based upon multi-resolution, spectro-temporal analysis and when fused with more traditional parameterisations, the new features not only generally outperform Mel-frequency cepstral coefficients, but also are shown to be complementary when fusing systems at score level. Finally, the joint operation of UV and ASV greatly decreases false acceptances for unmatched text trials. Héctor Delgado, Massimiliano Todisco, Md. Sahidullah, Achintya Kumar Sarkar, Nicholas W. D. Evans, Tomi Kinnunen, Zheng-Hua Tan |
SLT | 7 |
| 2016 | Speech enhancement using Long Short-Term Memory based recurrent Neural Networks for noise robust Speaker VerificationabstractIn this paper we propose to use a state-of-the-art Deep Recurrent Neural Network (DRNN) based Speech Enhancement (SE) algorithm for noise robust Speaker Verification (SV). Specifically, we study the performance of an i-vector based SV system, when tested in noisy conditions using a DRNN based SE front-end utilizing a Long Short-Term Memory (LSTM) architecture. We make comparisons to systems using a Non-negative Matrix Factorization (NMF) based front-end, and a Short-Time Spectral Amplitude Minimum Mean Square Error (STSA-MMSE) based front-end, respectively. We show in simulation experiments that a male-speaker and text-independent DRNN based SE front-end, without specific a priori knowledge about the noise type outperforms a text, noise type and speaker dependent NMF based front-end as well as a STSA-MMSE based front-end in terms of Equal Error Rates for a large range of noise types and signal to noise ratios on the RSR2015 speech corpus. Morten Kolbæk, Zheng-Hua Tan, Jesper Jensen 0001 |
SLT | 2 |
| 2016 | Feature selection for neutral vector in EEG signal classification
Zhanyu Ma, Zheng-Hua Tan, Jun Guo 0002 |
Neurocomputing | 2 |
| 2016 | AMORE: design and implementation of a commercial-strength parallel hybrid movie recommendation engine
Ioannis T. Christou, Emmanouil Amolochitis, Zheng-Hua Tan |
Knowl. Inf. Syst. | 3 |
| 2016 | Predicting the Intelligibility of Noisy and Nonlinearly Processed Binaural SpeechabstractObjective speech intelligibility measures are gaining popularity in the development of speech enhancement algorithms and speech processing devices such as hearing aids. Such devices may process the input signals nonlinearly and modify the binaural cues presented to the user. We propose a method for predicting the intelligibility of noisy and nonlinearly processed binaural speech. This prediction is based on the noisy and processed signal as well as a clean speech reference signal. The method is obtained by extending a modified version of the short-time objective intelligibility (STOI) measure with a modified equalization-cancellation (EC) stage. We evaluate the performance of the method by comparing the predictions with measured intelligibility from four listening experiments. These comparisons indicate that the proposed measure can provide accurate predictions of (1) the intelligibility of diotic speech with an accuracy similar to that of the original STOI measure, (2) speech reception thresholds (SRTs) in conditions with a frontal target speaker and a single interferer in the horizontal plane, (3) SRTs in conditions with a frontal target and a single interferer when ideal time frequency segregation (ITFS) is applied to the left and right ears separately, and (4) the advantage of two-microphone beamforming as applied in state-of-the-art hearing aids. A MATLAB implementation of the proposed measure is available online1. Asger Heidemann Andersen, Jan Mark de Haan, Zheng-Hua Tan, Jesper Jensen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Total Variability Modeling Using Source-Specific PriorsabstractIn total variability modeling, variable length speech utterances are mapped to fixed low-dimensional i-vectors. Central to computing the total variability matrix and i-vector extraction, is the computation of the posterior distribution for a latent variable conditioned on an observed feature sequence of an utterance. In both cases the prior for the latent variable is assumed to be non-informative, since for homogeneous datasets there is no gain in generality in using an informative prior. This work shows in the heterogeneous case, that using informative priors for computing the posterior, can lead to favorable results. We focus on modeling the priors using minimum divergence criterion or factor analysis techniques. Tests on the NIST 2008 and 2010 Speaker Recognition Evaluation (SRE) dataset show that our proposed method beats four baselines: For i-vector extraction using an already trained matrix, for the short2-short3 task in SRE'08, five out of eight female and four out of eight male common conditions, were improved. For the core-extended task in SRE'10, four out of nine female and six out of nine male common conditions were improved. When incorporating prior information into the training of the T matrix itself, the proposed method beats the baselines for six out of eight female and five out of eight male common conditions, for SRE'08, and five and six out of nine conditions, for the male and female case, respectively, for SRE'10. Tests using factor analysis for estimating priors show that two priors do not offer much improvement, but in the case of three separate priors (sparse data), considerable improvements were gained. Sven Ewan Shepstone, Kong-Aik Lee, Haizhou Li 0001, Zheng-Hua Tan, Søren Holdt Jensen |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2015 | Maximum likelihood approach to "informed" Sound Source Localization for Hearing Aid applicationsabstractMost state-of-the-art Sound Source Localization (SSL) algorithms have been proposed for applications which are “uninformed” about the target sound content; however, utilizing a wireless microphone worn by a target talker, enables recent Hearing Aid Systems (HASs) to access to an almost noise-free sound signal of the target talker at the HAS via the wireless connection. Therefore, in this paper, we propose a maximum likelihood (ML) approach, which we call MLSSL, to estimate the Direction of Arrival (DoA) of the target signal given access to the target signal content. Compared with other “informed” SSL algorithms which use binaural microphones for localization, MLSSL performs better using signals of one or more microphones placed on just one ear, thereby reducing the wireless transmission overhead of binaural hearing aids. More specifically, when the target location confined to the front-horizontal plane, MLSSL shows an average absolute DoA estimation error of 5 degrees at SNR of -5 dB in a large-crowd noise and non-reverberant situation. Moreover, MLSSL suffers less from front-back confusions compared with the recent approaches. Mojtaba Farmani, Michael Syskind Pedersen, Zheng-Hua Tan, Jesper Jensen 0001 |
ICASSP | 3 |
| 2015 | On the influence of microphone array geometry on HRTF-based Sound Source LocalizationabstractThe direction dependence of Head Related Transfer Functions (HRTFs) forms the basis for HRTF-based Sound Source Localization (SSL) algorithms. In this paper, we show how spectral similarities of the HRTFs of different directions in the horizontal plane influence performance of HRTF-based SSL algorithms; the more similar the HRTFs of different angles to the HRTF of the target angle, the worse the performance. However, we also show how the microphone array geometry can assist in differentiating between the HRTFs of the different angles, thereby improving performance of HRTF-based SSL algorithms. Furthermore, to demonstrate the analysis results, we show the impact of HRTFs similarities and microphone array geometry on an exemplary HRTF-based SSL algorithm, called MLSSL. This algorithm is well-suited for this purpose as it allows to estimate the Direction-of-Arrival (DoA) of the target sound using any number of microphones and any geometries of the microphone array around the head. Mojtaba Farmani, Michael Syskind Pedersen, Zheng-Hua Tan, Jesper Jensen 0001 |
ICASSP | 3 |
| 2015 | Source-specific informative prior for i-vector extractionabstractAn i-vector is a low-dimensional fixed-length representation of a variable-length speech utterance, and is defined as the posterior mean of a latent variable conditioned on the observed feature sequence of an utterance. The assumption is that the prior for the latent variable is non-informative, since for homogeneous datasets there is no gain in generality in using an informative prior. This work shows that extracting i-vectors for a heterogeneous dataset, containing speech samples recorded from multiple sources, using informative priors instead is applicable, and leads to favorable results. Tests carried out on the NIST 2008 and 2010 Speaker Recognition Evaluation (SRE) dataset show that our proposed method beats three baselines: For the short2-short3 core-task in SRE'08, for the female and male cases, five and six respectively, out of eight common conditions were beaten, and for the core-core task in SRE'10, for both genders, five out of nine common conditions were beaten. Sven Ewan Shepstone, Kong-Aik Lee, Haizhou Li 0001, Zheng-Hua Tan, Søren Holdt Jensen |
ICASSP | 4 |
| 2015 | A feature subtraction method for image based kinship verification under uncontrolled environmentsabstractThe most fundamental problem of local feature based kinship verification methods is that a local feature can capture the variations of environmental conditions and the differences between two persons having a kin relation, which can significantly decrease the performance. To address this problem, we propose a feature subtraction method to remove the kinship unrelated part from the local feature through a linear function of which only one parameter, namely a subtraction matrix, needs to be inferred from training data. This is done by using a gradient descent method to simultaneously minimize the feature distance between face image pairs with kinship and maximize the distance between non-kinship pairs. Based on the subtracted feature, the verification is realized through a simple Gaussian based distance comparison method. Experiments on two public databases show that the feature subtraction method outperforms or is comparable to state-of-the-art kinship verification methods. Xiaodong Duan, Zheng-Hua Tan |
ICIP | 2 |
| 2015 | Local feature learning for face recognition under varying posesabstractIn this paper, we present a local feature learning method for face recognition to deal with varying poses. As opposed to the commonly used approaches of recovering frontal face images from profile views, the proposed method extracts the subject related part from a local feature by removing the pose related part in it on the basis of a pose feature. The method has a closed-form solution, hence being time efficient. For performance evaluation, cross pose face recognition experiments are conducted on two public face recognition databases FERET and FEI. The proposed method shows a significant recognition improvement under varying poses over general local feature approaches and outperforms or is comparable with related state-of-the-art pose invariant face recognition approaches. Xiaodong Duan, Zheng-Hua Tan |
ICIP | 2 |
| 2015 | A binaural short time objective intelligibility measure for noisy and enhanced speechabstractObjective intelligibility measures are increasingly being used to assess the performance of speech processing algorithms, e.g. for hearing aids. It has been shown that the short time objective intelligibility (STOI) measure yields good results in this respect. In this paper we propose a binaural extension of the STOI measure, which predicts binaural advantage using a modified equalization cancellation (EC) stage. The proposed method is evaluated for a range of acoustic conditions. Firstly, the method is able to predict the advantage of spatial separation between a speech target and a speech shaped noise (SSN) interferer. Secondly, the method yields results comparable to the monaural STOI measure when presented with noisy speech processed by ideal time-frequency segregation (ITFS). Finally, the method also performs well when presented with a selection of different acoustic conditions combined with beamforming as used in hearing aids. Asger Heidemann Andersen, Jan Mark de Haan, Zheng-Hua Tan, Jesper Jensen 0001 |
INTERSPEECH | 3 |
| 2015 | Comparison of forced-alignment speech recognition and humans for generating reference VADabstractThis present paper aims to answer the question whether forced-alignment speech recognition can be used as an alternative to humans in generating reference Voice Activity Detection (VAD) transcriptions. An investigation of the level of agreement between automatic/manual VAD transcriptions and the reference ones produced by a human expert was carried out. Thereafter, statistical analysis was employed on the automatically produced and the collected manual transcriptions. Experimental results confirmed that forced-alignment speech recognition can provide accurate and consistent VAD labels. Ivan Kraljevski, Zheng-Hua Tan, Maria Paola Bissiri |
INTERSPEECH | 2 |
| 2015 | A heuristic approach for a social robot to navigate to a person based on audio and range informationabstractThe use of social robots for elderly care is becoming ever more relevant, thus introducing new challenges which need to be solved to achieve acceptable performance. One fundamental task for a social robot is to move to the person of interest in order to start interacting or perform a service. In this paper we address the task of a robot having to navigate to a possibly occluded person, which needs assistance, based only on audio and range information. Our approach is based on forming a heuristic cost function which is based on combining two audio features, and then moving to the optimum position indicated by this cost function after every interaction with the person. The method is compared to a greedy approach in 20 different tasks using a loud speaker playing at approximately 60dB sound pressure level (SPL) to mimic a human speaker, and the proposed method shows superior performance. A second experiment with an increase of 13dB SPL of the loud speaker is conducted and the proposed method is able to handle this. Nicolai Bæk Thomsen, Zheng-Hua Tan, Børge Lindberg, Søren Holdt Jensen |
IROS | 2 |
| 2015 | Im2Sketch: Sketch generation by unconflicted perceptual grouping
Yonggang Qi, Jun Guo 0002, Yi-Zhe Song, Tao Xiang 0002, Honggang Zhang 0002, Zheng-Hua Tan |
Neurocomputing | 6 |
| 2015 | Minimum Mean-Square Error Estimation of Mel-Frequency Cepstral Features-A Theoretically Consistent ApproachabstractIn this work, we consider the problem of feature enhancement for noise-robust automatic speech recognition (ASR). We propose a method for minimum mean-square error (MMSE) estimation of mel-frequency cepstral features, which is based on a minimum number of well-established, theoretically consistent statistical assumptions. More specifically, the method belongs to the class of methods relying on the statistical framework proposed in Ephraim and Malah's original work (“Speech enhancement using a minimum mean-square error short-time spectral amplitude estimator,” IEEE Trans. Acoust., Speech, Signal Process., vol. ASSP-32, no. 6, 1984). The method is general in that it allows MMSE estimation of mel-frequency cepstral coefficients (MFCC's), cepstral-mean subtracted (CMS-) MFCC's, autoregressive-moving-average (ARMA)-filtered CMS-MFCC's, velocity, and acceleration coefficients. In addition, the method is easily modified to take into account other compressive non-linearities than the logarithm traditionally used for MFCC computation. In terms of MFCC estimation performance, as measured by MFCC mean-square error, the proposed method shows performance which is identical to or better than other state-of-the-art methods. In terms of ASR performance, no statistical difference could be found between the proposed method and the state-of-the-art methods. We conclude that existing state-of-the-art MFCC feature enhancement algorithms within this class of algorithms, while theoretically suboptimal or based on theoretically inconsistent assumptions, perform close to optimally in the MMSE sense. Jesper Jensen 0001, Zheng-Hua Tan |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Using Audio-Derived Affective Offset to Enhance TV RecommendationabstractThis paper introduces the concept of affective offset, which is the difference between a user's perceived affective state and the affective annotation of the content they wish to see. We show how this affective offset can be used within a framework for providing recommendations for TV programs. First a user's mood profile is determined using 12-class audio-based emotion classifications . An initial TV content item is then displayed to the user based on the extracted mood profile. The user has the option to either accept the recommendation, or to critique the item once or several times, by navigating the emotion space to request an alternative match. The final match is then compared to the initial match, in terms of the difference in the items' affective parameterization . This offset is then utilized in future recommendation sessions. The system was evaluated by eliciting three different moods in 22 separate users and examining the influence of applying affective offset to the users' sessions. Results show that, in the case when affective offset was applied, better user satisfaction was achieved: the average ratings went from 7.80 up to 8.65, with an average decrease in the number of critiquing cycles which went from 29.53 down to 14.39. Sven Ewan Shepstone, Zheng-Hua Tan, Søren Holdt Jensen |
IEEE Trans. Multim. | 2 |
| 2013 | Developing a speaker identification system for the DARPA RATS projectabstractThis paper describes the speaker identification (SID) system developed by the Patrol team for the first phase of the DARPA RATS (Robust Automatic Transcription of Speech) program, which seeks to advance state of the art detection capabilities on audio from highly degraded communication channels. We present results using multiple SID systems differing mainly in the algorithm used for voice activity detection (VAD) and feature extraction. We show that (a) unsupervised VAD performs as well supervised methods in terms of downstream SID performance, (b) noise-robust feature extraction methods such as CFCCs out-perform MFCC front-ends on noisy audio, and (c) fusion of multiple systems provides 24% relative improvement in EER compared to the single best system when using a novel SVM-based fusion algorithm that uses side information such as gender, language, and channel id. Oldrich Plchot, Spyridon Matsoukas, Pavel Matejka, Najim Dehak, Jeff Z. Ma, Sandro Cumani, Ondrej Glembek, Hynek Hermansky, Sri Harish Reddy Mallidi, Nima Mesgarani, Richard M. Schwartz, Mehdi Soufifar, Zheng-Hua Tan, Samuel Thomas 0001, Bing Zhang 0004, Xinhui Zhou |
ICASSP | 13 |
| 2013 | Demographic recommendation by means of group profile elicitation using speaker age and gender recognitionabstractIn this paper we show a new method of using automatic age and gender recognition to recommend a sequence of multimedia items to a home TV audience comprising multiple viewers. Instead of relying on explicitly provided demographic data for each user, we define an audio-based demographic group profile that captures the age and gender for all members of the audience. A 7-class age and gender classifier employing a fusion of acoustic and prosodic features determines the probability of each speaker belonging to each class. The information for all speakers is then combined to form the group profile, which itself is the input to a recommender system. The recommender system finds the content items whose demographics best match the group profile. We tested the effectiveness of the system for several typical home audience configurations. In a survey, users were given a configuration and asked to rate a set of advertisements on how well each advertisement matched the configuration. Unbeknown to the subjects, half of the adverts were recommended using the derived audio demographics and the other half were randomly chosen. The recommended adverts received a significantly higher median rating of 7.75, as opposed to 4.25 for the randomly selected adverts. Sven Ewan Shepstone, Zheng-Hua Tan, Søren Holdt Jensen |
INTERSPEECH | 2 |
| 2013 | Perceptual grouping via untangling Gestalt principlesabstractGestalt principles, a set of conjoining rules derived from human visual studies, have been known to play an important role in computer vision. Many applications such as image segmentation, contour grouping and scene understanding often rely on such rules to work. However, the problem of Gestalt confliction, i.e., the relative importance of each rule compared with another, remains unsolved. In this paper, we investigate the problem of perceptual grouping by quantifying the confliction among three commonly used rules: similarity, continuity and proximity. More specifically, we propose to quantify the importance of Gestalt rules by solving a learning to rank problem, and formulate a multi-label graph-cuts algorithm to group image primitives while taking into account the learned Gestalt confliction. Our experiment results confirm the existence of Gestalt confliction in perceptual grouping and demonstrate an improved performance when such a confliction is accounted for via the proposed grouping algorithm. Finally, a novel cross domain image classification method is proposed by exploiting perceptual grouping as representation. Yonggang Qi, Jun Guo 0002, Yi Li 0004, Honggang Zhang 0002, Tao Xiang 0002, Yi-Zhe Song, Zheng-Hua Tan |
VCIP | 7 |
| 2013 | A heuristic hierarchical scheme for academic search and retrieval
Emmanouil Amolochitis, Ioannis T. Christou, Zheng-Hua Tan, Ramjee Prasad |
Inf. Process. Manag. | 3 |
| 2012 | PubSearch - A Hierarchical Heuristic Scheme for Ranking Academic Search Results
Emmanouil Amolochitis, Ioannis T. Christou, Zheng-Hua Tan |
ICPRAM (2) | 3 |
| 2012 | A Joint Approach for Single-Channel Speaker Identification and Speech SeparationabstractIn this paper, we present a novel system for joint speaker identification and speech separation. For speaker identification a single-channel speaker identification algorithm is proposed which provides an estimate of signal-to-signal ratio (SSR) as a by-product. For speech separation, we propose a sinusoidal model-based algorithm. The speech separation algorithm consists of a double-talk/single-talk detector followed by a minimum mean square error estimator of sinusoidal parameters for finding optimal codevectors from pre-trained speaker codebooks. In evaluating the proposed system, we start from a situation where we have prior information of codebook indices, speaker identities and SSR-level, and then, by relaxing these assumptions one by one, we demonstrate the efficiency of the proposed fully blind system. In contrast to previous studies that mostly focus on automatic speech recognition (ASR) accuracy, here, we report the objective and subjective results as well. The results show that the proposed system performs as well as the best of the state-of-the-art in terms of perceived quality while its performance in terms of speaker identification and automatic speech recognition results are generally lower. It outperforms the state-of-the-art in terms of intelligibility showing that the ASR results are not conclusive. The proposed method achieves on average, 52.3% ASR accuracy, 41.2 points in MUSHRA and 85.9% in speech intelligibility. Pejman Mowlaee, Rahim Saeidi, Mads Græsbøll Christensen, Zheng-Hua Tan, Tomi Kinnunen, Pasi Fränti, Søren Holdt Jensen |
IEEE Trans. Speech Audio Process. | 4 |
| 2011 | Sinusoidal Approach for the Single-Channel Speech Separation and Recognition ChallengeabstractMost of the single-channel speech separation (SCSS) systems use the short-time Fourier transform as their parametric features. Recent studies have shown that employing sinusoidal features for the SCSS application results in a high perceived speech quality. In this paper, we make a systematic study on automatic speech recognition results for a SCSS system that uses sinusoidal features composed of amplitude and frequency. We compare the speech recognition results with those already reported by other participants in the single-channel speech separation and recognition challenge. Our results show that a newly proposed system achieves an overall recognition accuracy of 52.3%, ranges at the median over all other participants in the challenge. Index Terms: sinusoidal modeling, single-channel speech separation and recognition challenge. Pejman Mowlaee, Rahim Saeidi, Zheng-Hua Tan, Mads Græsbøll Christensen, Tomi Kinnunen, Pasi Fränti, Søren Holdt Jensen |
INTERSPEECH | 3 |
| 2011 | Multi-Sensor Voice Activity Detection Based on Multiple Observation Hypothesis Testing
Theodore Petsatodis, Fotios Talantzis, Christos Boukis, Zheng-Hua Tan, Ramjee Prasad |
INTERSPEECH | 4 |
| 2011 | Convex Combination of Multiple Statistical Models With Application to VADabstractThis paper proposes a robust voice activity detector (VAD) based on the observation that the distribution of speech captured with far-field microphones is highly varying, depending on the noise and reverberation conditions. The proposed VAD employs a convex combination scheme comprising three statistical distributions - a Gaussian, a Laplacian, and a two-sided Gamma - to effectively model captured speech. This scheme shows increased ability to adapt to dynamic acoustic environments. The contribution of each distribution to this convex combination is automatically adjusted based on the statistical characteristics of the instantaneous audio input. To further improve the performance of the system, an adaptive threshold is introduced, while a decision-smoothing scheme caters to the intra-frame correlation of speech signals. Extensive experiments under realistic scenarios support the proposed approach of combining several models for increased adaptation and performance. Theodore Petsatodis, Christos Boukis, Fotios Talantzis, Zheng-Hua Tan, Ramjee Prasad |
IEEE Trans. Speech Audio Process. | 4 |
| 2010 | Joint single-channel speech separation and speaker identificationabstractIn this paper, we propose a closed loop system to improve the performance of single-channel speech separation in a speaker independent scenario. The system is composed of two interconnected blocks: a separation block and a speaker identification block. The improvement is accomplished by incorporating the speaker identities found by the speaker identification block as additional information for the separation block, which converts the speaker-independent separation problem to a speaker-dependent one where the speaker codebooks are known. Simulation results show that the closed loop system enhances the quality of the separated output signals. To assess the improvements, the results are reported in terms of PESQ for both target and masked signals. Pejman Mowlaee, Rahim Saeidi, Zheng-Hua Tan, Mads Græsbøll Christensen, Pasi Fränti, Søren Holdt Jensen |
ICASSP | 3 |
| 2010 | Signal-to-Signal Ratio Independent Speaker Identification for Co-channel Speech SignalsabstractIn this paper, we consider speaker identification for the co-channel scenario in which speech mixture from speakers is recorded by one microphone only. The goal is to identify both of the speakers from their mixed signal. High recognition accuracies have already been reported when an accurately estimated signal-to-signal ratio (SSR) is available. In this paper, we approach the problem without estimating SSR. We show that a simple method based on fusion of adapted Gaussian mixture models and Kullback-Leibler divergence calculated between models, achieves an accuracy of 97% and 93% when the two target speakers enlisted as three and two most probable speakers, respectively. Rahim Saeidi, Pejman Mowlaee, Tomi Kinnunen, Zheng-Hua Tan, Mads Græsbøll Christensen, Søren Holdt Jensen, Pasi Fränti |
ICPR | 4 |
| 2010 | Improving monaural speaker identification by double-talk detectionabstractThis paper describes a novel approach to improve monoaural speaker identification where two speakers are present in a single-microphone recording. The goal is to identify both of the underlying speakers in the given mixture. The proposed approach is composed of a double-talk detector (DTD) as a preprocessor and speaker identification back-end. We demonstrate that including the double-talk detector improves the speaker identification accuracy. Experiments on GRID corpus show that including the DTD improves average recognition accuracy from 96.53% to 97.43%. Rahim Saeidi, Pejman Mowlaee, Tomi Kinnunen, Zheng-Hua Tan, Mads Græsbøll Christensen, Søren Holdt Jensen, Pasi Fränti |
INTERSPEECH | 4 |
| 2009 | A system for detecting miscues in dyslexic read speechabstractWhile miscue detection in general is a well explored research field little attention has so far been paid to miscue detection in dyslexic read speech. This domain differs substantially from the domains that are commonly researched, as for example dyslexic read speech includes frequent regressions and long pauses between words. A system detecting miscues in dyslexic read speech is presented. It includes an ASR component employing a forced-alignment like grammar adjusted for dyslexic input and uses the GOP score and phone duration to accept or reject the read words. Experimental results show that the system detects miscues at a false alarm rate of 5.3% and a miscue detection rate of 40.1%. These results are worse than current state of the art reading tutors perhaps indicating that dyslexic read speech is a challenge to handle. Morten Højfeldt Rasmussen, Zheng-Hua Tan, Børge Lindberg, Søren Holdt Jensen |
INTERSPEECH | 2 |
| 2009 | High-accuracy, low-complexity voice activity detection based on a posteriori SNR weighted energyabstractThis paper presents a voice activity detection (VAD) method using the measurement of a posteriori signal-to-noise ratio (SNR) weighted energy.The motivations are manifold: 1) the difference in frame-to-frame energy provides a great discrimination for speech signals, 2) speech segments, besides their characteristics, are accounted also on their reliability e.g.measured by SNR, 3) the a posteriori SNR for noise-only segments will theoretically equal to 0 dB, being ideal for VAD, and 4) both energy and a posteriori SNR are easy to estimate, resulting in a low complexity.The method is experimentally shown to be superior to a number of referenced methods and standards. Zheng-Hua Tan, Børge Lindberg |
INTERSPEECH | 1 |
| 2008 | A posteriori SNR weighted energy based variable frame rate analysis for speech recognitionabstractThis paper presents a variable frame rate (VFR) analysis method that uses an a posteriori signal-to-noise ratio (SNR) weighted energy distance for frame selection. The novelty of the method consists in the use of energy distance (instead of cepstral distance) to make it computationally efficient and the use of SNR weighting to emphasize the reliable regions in speech signals. The VFR method is applied to speech recognition in two scenarios. First, it is used for improving speech recognition performance in noisy environments. Secondly, the method is used for source coding in distributed speech recognition where the target bit rate is met by adjusting the frame rate, yielding a scalable coding scheme. Prior to recognition in the server, frames are repeated so that the original frame rate is restored. Very encouraging results are obtained for both noise robustness and source coding. Index Terms: speech recognition, speech analysis, variable frame rate, noise robustness, source coding Zheng-Hua Tan, Børge Lindberg |
INTERSPEECH | 1 |
| 2008 | Robust Speech Recognition by Nonlocal Means Denoising ProcessingabstractThe nonlocal means (NL-means) algorithm recently proposed for image denoising has proved highly effective for removing additive noise while to a large extent maintaining image details. The algorithm performs denoising by averaging each pixel with other pixels that have similar characteristics in the image. This letter considers the real and imaginary parts of complex speech spectrogram each as a separate image and presents a modified NL-means algorithm to them for denoising to improve the noise robustness of speech recognition. Recognition results on a noisy speech database show that the proposed method is superior to classical methods such as spectral subtraction. Haitian Xu, Zheng-Hua Tan, Paul Dalsgaard, Børge Lindberg |
IEEE Signal Process. Lett. | 2 |
| 2007 | Exploiting Temporal Correlation of Speech for Error Robust and Bandwidth Flexible Distributed Speech RecognitionabstractIn this paper, the temporal correlation of speech is exploited in front-end feature extraction, client-based error recovery, and server-based error concealment (EC) for distributed speech recognition. First, the paper investigates a half frame rate (HFR) front-end that uses double frame shifting at the client side. At the server side, each HFR feature vector is duplicated to construct a full frame rate (FFR) feature sequence. This HFR front-end gives comparable performance to the FFR front-end but contains only half the FFR features. Second, different arrangements of the other half of the FFR features creates a set of error recovery techniques encompassing multiple description coding and interleaving schemes where interleaving has the advantage of not introducing a delay when there are no transmission errors. Third, a subvector-based EC technique is presented where error detection and concealment is conducted at the subvector level as opposed to conventional techniques where an entire vector is replaced even though only a single bit error occurs. The subvector EC is further combined with weighted Viterbi decoding. Encouraging recognition results are observed for the proposed techniques. Lastly, to understand the effects of applying various EC techniques, this paper introduces three approaches consisting of speech feature, dynamic programming distance, and hidden Markov model state duration comparison. Zheng-Hua Tan, Paul Dalsgaard, Børge Lindberg |
IEEE Trans. Speech Audio Process. | 1 |
| 2007 | Noise Condition-Dependent Training Based on Noise Classification and SNR EstimationabstractCondition-dependent training strategy divides a training database into a number of clusters, each corresponding to a noise condition and subsequently trains a hidden Markov model (HMM) set for each cluster. This paper investigates and compares a number of condition-dependent training strategies in order to achieve a better understanding of the effects on automatic speech recogntion (ASR) performance as caused by a splitting of the training databases. Also, the relationship between mismatches in signal-to-noise ratio (SNR) is analyzed. The results show that a splitting of the training material in terms of both noise type and SNR value is advantageous compared to previously used methods, and that training of only a limited number of HMM sets is sufficient for each noise type for robustly handling of SNR mismatches. This leads to the introduction of an SNR and noise classification-based training strategy (SNT-SNC). Better ASR performance is obtained on test material containing data from known noise types as compared to either multicondition training or noise-type dependent training strategies. The computational complexity of the SNT-SNC framework is kept low by choosing only one HMM set for recognition. The HMM set is chosen on the basis of results from noise classification and SNR value estimations. However, compared to other strategies, the SNT-SNC framework shows lower performance for unknown noise types. This problem is partly overcome by introducing a number of model and feature domain techniques. Experiments using both artificially corrupted and real-world noisy speech databases are conducted and demonstrate the effectiveness of these methods. Haitian Xu, Paul Dalsgaard, Zheng-Hua Tan, Børge Lindberg |
IEEE Trans. Speech Audio Process. | 3 |
| 2006 | Robust Speech Recognition From Noise-Type Based Feature Compensation and Model Interpolation in a Multiple Model FrameworkabstractCompared to multi-condition training (MTR), condition-dependent training generates multiple acoustic hidden Markov model sets each identified by a noisy environment and is known to perform substantially better for known noise types (included in training) while worse for unknown (untrained) noise types. This paper attempts to bridge the performance gap between known and unknown noise types by introducing a Minimum Mean-Square Error (MMSE) noise-type based compensation algorithm. On the basis of a modified Vector Taylor Series and the measurement of feature reliability as well as noise similarity, the MMSE estimation adapts the test features corrupted by the unknown noise type to the corresponding features corrupted by the known noise type. This method significantly improves the recognition performance for unknown noise types while maintaining the good performance for known noise types. Furthermore, in order to benefit directly from MTR, a model interpolation strategy is investigated which combines the MTR and the condition-dependent model sets. Both good performance and low computational cost are achieved by only interpolating the mixtures of each condition-dependent model state with the least weighted mixture in the corresponding MTR model state. The overall system gives promising results. Haitian Xu, Zheng-Hua Tan, Paul Dalsgaard, Børge Lindberg |
ICASSP (1) | 2 |
| 2006 | Robust speech recognition over mobile networks using combined weighted viterbi decoding and subvector based error concealmentabstractRobustness against transmission errors is one of the primary barriers to the widespread application of automatic speech recognition (ASR) in mobile communications. We have previously proposed a subvector based error concealment (EC) method that conducts error detection and mitigation in the feature-domain at the subvector level. This paper presents a weighted Viterbi decoding (WVD) algorithm that works in the model domain for counteracting unreliable features generated by the subvector based EC. The reliability of each feature is estimated during the process of subvector based EC and is used by the WVD for modifying the observation probability of the feature. Recognition experiments are conducted on the Aurora 2 database corrupted by GSM error pattern EP3. Combining the WVD and the subvector EC achieves 70 % and 24% performance improvement as compared to the ETSI-DSR standard and the subvector based EC, respectively. Index Terms: distributed speech recognition, error concealment, split vector quantization, weighted Viterbi Zheng-Hua Tan, Paul Dalsgaard, Børge Lindberg |
INTERSPEECH | 1 |
| 2006 | Fuzzy Metagraph and Its Combination with the Indexing Approach in Rule-Based SystemsabstractThis paper presents a graph-theoretic construct called a fuzzy metagraph (FM) with the capability of describing the relationships between sets of fuzzy elements instead of only single fuzzy elements. The algebraic structure of FM and its properties are extensively investigated. Subsequently, the FM construct is applied to rule-based systems. First, we propose FM-based knowledge representation in both graphic and algebraic format. The representation is capable of identifying dependencies across compound propositions in the rules. In the algebraic representation, the FM closure matrix is considered a precompiled rule base enabling efficient query processing. An iterative approach is presented to facilitate the construction and expansion of the FM closure matrix, which is a key for real-world applications. Next, we introduce the concept of indexing, which was originally developed for information retrieval (IR), to enable an immediate extraction of relevant entries from the FM closure matrix. The indexing approach is applied in combination with the FM closure matrix. Based on the combination, corresponding inference mechanisms are introduced to achieve instant acquisition of relevant rules over a large collection of rules. The application in rule-based systems indicates that the combination of FM and IR techniques offers advantages for the mathematical analysis of systems. Zheng-Hua Tan |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2005 | Robust speech recognition in ubiquitous networking and context-aware computingabstractThe introduction of ubiquitous computing and networking has fostered automatic speech recognition (ASR) systems of a distributed nature. The major challenge in deploying ubiquitous ASR is that the operating environments may change rapidly leaving the ASR system very vulnerable. This paper deals with the concept of making ASR systems context-aware with the aim of improving robustness against varying conditions such as dynamic network constraints and environmental noise. To fully benefit from a variety of networks with different characteristics, a number of distributed speech recognition (DSR) schemes are presented each of which is applicable to a specific network context. To increase ASR system robustness in varying environmental noise context, a multiple-model framework for noise-robust ASR is presented where multiple HMM model sets are trained, one for each noise type and each specific signal-to-noise ratio (SNR) that characterise the noise context. Experimental results show that the performance of ASR is largely improved by exploiting the context information. 1. Zheng-Hua Tan, Paul Dalsgaard, Børge Lindberg, Haitian Xu |
INTERSPEECH | 1 |
| 2005 | Robust speech recognition based on noise and SNR classification - a multiple-model frameworkabstractThis paper presents a multiple-model framework for noise-robust speech recognition. In this framework, multiple HMM model sets are trained- each identified by a noise type and a specific Signal-to-Noise Ratio (SNR) value. This, however, does not increase the computational complexity of the recognition process since only one model set is selected according to the noise classification and SNR estimation. The optimal number of model sets is first identified on the basis of the Aurora 2 database. With only three model sets for each noise type, the framework shows superior performance to Multi-style TRaining (MTR) when testing on known noise types but lower performance on unknown noise types. To overcome this drawback, a modified Jacobian method is proposed to adapt the selected HMM models to the test environment. Furthermore, given the fact that MTR often gives relatively stable performance for unknown noise types, a combined technique is applied in which interpolation between the MTR and the adapted models is performed. This combined technique gives more than 24 % performance improvement as compared to MTR. 1. Haitian Xu, Zheng-Hua Tan, Paul Dalsgaard, Børge Lindberg |
INTERSPEECH | 2 |
| 2005 | Adaptive Multi-Frame-Rate Scheme for Distributed Speech Recognition Based on a Half Frame-Rate Front-EndabstractIn this paper a half frame-rate (HFR) front-end is investigated for distributed speech recognition (DSR). The work is inspired from the need for low bit-rate and is justified by the redundancies known to exist in full frame-rate (FFR) features. At the client-side in the DSR architecture, implementation of the HFR is carried out by using double frame shifting as compared to the FFR resulting in the achievement of half the bit rate. At the server-side, each HFR feature vector is repeated once to construct the FFR features and no changes are therefore required in the recognition back-end. It is experimentally justified that the performance achieved by HFR is comparable to FFR and that repetition of each HFR feature vector is critical for the HFR front-end to maintain the performance. Motivated by the effectiveness of HFR, a number of additional FFR-based DSR schemes are further presented. Finally, this paper introduces an adaptive multi-frame-rate scheme in which the DSR system adapts to the characteristics of the transmission channel by switching between HFR and the FFR-based schemes. This multi-frame-rate scheme is found to be superior to the basic FFR Zheng-Hua Tan, Paul Dalsgaard, Børge Lindberg |
MMSP | 1 |
| 2005 | Automatic speech recognition over error-prone wireless networks
Zheng-Hua Tan, Paul Dalsgaard, Børge Lindberg |
Speech Commun. | 1 |
| 2004 | A subvector-based error concealment algorithm for speech recognition over mobile networksabstractConventional error concealment (EC) algorithms for distributed speech recognition (DSR) share a common characteristic namely the fact of conducting EC at the vector (or frame) level. This strategy, however, fails to effectively exploit the error-free fraction left within erroneous vectors where a substantial number of subvectors often are error-free. This paper proposes a novel EC approach for DSR encoded by split vector quantization (SVQ) where the detected erroneous vectors are submitted to a further analysis at the subvector level. Specifically, a data consistency test is applied to each erroneous vector to identify inconsistent subvectors. Only inconsistent subvectors are replaced by their nearest neighbouring consistent subvectors whereas consistent subvectors are kept untouched. Experimental results demonstrate that the proposed algorithm in terms of recognition accuracy is superior to conventional EC methods having almost the same complexity and resource requirement. Zheng-Hua Tan, Paul Dalsgaard, Børge Lindberg |
ICASSP (1) | 1 |
| 2004 | On the integration of speech recognition into personal networksabstractMobile communication presents a number of challenges to speech technology such as the limited resources available in the terminals in addition to the bandwidth constraints and the errors occurring in transmissions over mobile networks. These challenges need to be solved before automatic speech recognition (ASR) is ready for widespread use in the context of personal communication environments. This paper gives an overview of the problems inherent in the recently developed network based ASR with an emphasis on the robustness issues that are highly influenced by network degradations. The paper further presents a number of transmission error protection and concealment schemes that are evaluated in a number of ASR experiments encompassing a range of typical real-environment transmission errors. 1. Zheng-Hua Tan, Paul Dalsgaard, Børge Lindberg |
INTERSPEECH | 1 |
| 2004 | Spectral subtraction with full-wave rectification and likelihood controlled instantaneous noise estimation for robust speech recognitionabstractIn standard Spectral Subtraction (SS), Half-Wave Rectification SS (HWR-SS) is normally applied to avoid negative values in the Power Spectral Density (PSD) that occur mainly due to inaccurate noise estimation caused by a Voice Activity Detector (VAD). In this paper analyses show that, given accurate noise estimation, the phase relationship between speech and noise becomes the dominant cause of the negative values. FullWave Rectification based SS (FWR-SS) combined with Instantaneous Noise Estimation (INE) is therefore proposed to be applied instead of VAD based HWR-SS as it is better capable of maintaining the speech information in those negative values. It is also shown in the paper that FWR-SS provides optimum orthogonality between the estimated noise and speech signals. The INE method proposed in this paper is Likelihood Controlled Instantaneous Noise Estimation (LCINE), which combines long-term statistical characteristics of noise resulting from a VAD with a method of short-term INE. The combination of FWR-SS and LCINE is computationally efficient and shows a 51% error rate reduction on the Aurora 2 database in comparison to the basic Aurora front-end provided by ETSI [1]. Haitian Xu, Zheng-Hua Tan, Paul Dalsgaard, Børge Lindberg |
INTERSPEECH | 2 |
| 2003 | OOV-detection and channel error protection for distributed speech recognition over wireless networksabstractThis paper presents research on two aspects of distributed speech recognition (DSR) in the presence of channel transmission errors in wireless network environments. The first is on experiments with a frame-based channel error protection scheme, where in previous research we reported results from experiments using randomly distributed bit-errors. This paper presents results from experiments using three additional, more realistic error distributions: burst-like packet loss, GSM error patterns and UMTS statistics. The second is on exploiting the knowledge about channel transmission errors for the purpose of optimising the Out-of-Vocabulary (OOV) detection. Transmission errors influence the acoustic likelihood, and therefore affect the optimal threshold setting for discrimination between In-Vocabulary (IV) words and OOV words. An OOV-detection method is proposed in which the estimated Frame-Error-Rate (FER) is used to adjust the discrimination threshold. Results from experiments are reported over a range of transmission errors. Zheng-Hua Tan, Paul Dalsgaard, Børge Lindberg |
ICASSP (1) | 1 |
| 2002 | Channel error protection scheme for distributed speech recognitionabstractThis paper describes ongoing research preparing for the widespread deployment of spoken language processing in networks encompassing wired and wireless transmission channels. The paper gives a brief overview of the standardized bit-error protection scheme aimed at minimising channel transmission errors and used within the distributed speech recognition (DSR) paradigm. Within the ETSI-DSR standard, two quantised mel-spectral frames – each of 10 ms duration- are grouped together and protected with a 4-bit Cyclic Redundancy Checking (CRC) forming a frame-pair. However, this causes the entire frame-pair erroneous if a one-bit error only occurs in the frame-pair packet. Over an error-prone transmission channel this format will cause severe problems. To overcome this, the paper presents a one-frame architecture in which a 4-bit CRC is calculated to protect each frame independently. This scheme results in that the overall probability of one frame in error is lower, or that an error occurring in one frame does not affect another frame. A number of simple recognition experiments have been conducted to verify the introduction of the one-frame CRC protection scheme for a number of simulated transmission channel bit-error rates (BER) ranging from 0 (no transmission channel involved) to 2٠10-2. Experimental results show that the one-frame protection scheme is more robust to channel errors although a slight increase in the error-protection overhead is needed. 1. Zheng-Hua Tan, Paul Dalsgaard |
INTERSPEECH | 1 |