Nicholas W. D. Evans

dblp:84/366 · DBLP profile ↗
← Back
110ranked-venue papers
7as first author
35since 2021 · last 2026
0000-0002-8459-1041ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 90 · 5 first-author · 25 since 2021Artificial intelligence and machine learning · 70 · 4 first-author · 25 since 2021Security and privacy · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 The third VoicePrivacy challenge: Preserving emotional expressiveness and linguistic content in voice anonymization
Natalia A. Tomashenko, Xiaoxiao Miao, Pierre Champion, Sarina Meyer, Michele Panariello, Xin Wang 0037, Nicholas W. D. Evans, Emmanuel Vincent 0001, Junichi Yamagishi, Massimiliano Todisco
Comput. Speech Lang.7
2026 ASVspoof 5: Design, collection and validation of resources for spoofing, deepfake, and adversarial attack detection using crowdsourced speech
abstract
ASVspoof 5 is the fifth edition in a series of challenges which promote the study of speech spoofing and deepfake attacks as well as the design of detection solutions. We introduce the ASVspoof 5 database which is generated in a crowdsourced fashion from data collected in diverse acoustic conditions (cf. studio-quality data for earlier ASVspoof databases) and from ∼ 2,000 speakers (cf. ∼ 100 earlier). The database contains attacks generated with 32 different algorithms, also crowdsourced, and optimised to varying degrees using new surrogate detection models. Among them are attacks generated with a mix of legacy and contemporary text-to-speech synthesis and voice conversion models, in addition to adversarial attacks which are incorporated for the first time. ASVspoof 5 protocols comprise seven speaker-disjoint partitions. They include two distinct partitions for the training of different sets of attack models, two more for the development and evaluation of surrogate detection models, and then three additional partitions which comprise the ASVspoof 5 training, development and evaluation sets. An auxiliary set of data collected from an additional 30k speakers can also be used to train speaker encoders for the implementation of attack algorithms. Also described herein is an experimental validation of the new ASVspoof 5 database using a set of automatic speaker verification and spoof/deepfake baseline detectors. With the exception of protocols and tools for the generation of spoofed/deepfake speech, the resources described in this paper, already used by participants of the ASVspoof 5 challenge in 2024, are now all freely available to the community.
Xin Wang 0037, Héctor Delgado, Hemlata Tak, Jee-Weon Jung, Hye-Jin Shim, Massimiliano Todisco, Ivan Kukanov, Xuechen Liu 0001, Md. Sahidullah, Tomi Kinnunen, Nicholas W. D. Evans, Kong-Aik Lee, Junichi Yamagishi, Myeonghun Jeong, Yongyi Zang, Soumi Maiti, Florian Lux, Nicolas Müller, Wangyou Zhang, Chengzhe Sun 0001, Shuwei Hou, Siwei Lyu, Sébastien Le Maguer, Hanjie Guo, Vishwanath Pratap Singh
Comput. Speech Lang.11
2025 The Text-to-speech in the Wild (TITW) Database
Jee-Weon Jung, Wangyou Zhang, Soumi Maiti, Yihan Wu 0008, Xin Wang 0037, Yuta Matsunaga, Seyun Um, Jinchuan Tian, Hye-Jin Shim, Nicholas W. D. Evans, Joon Son Chung, Shinnosuke Takamichi, Shinji Watanabe 0001
INTERSPEECH11
2024 Spoofing Attack Augmentation: Can Differently-Trained Attack Models Improve Generalisation?
abstract
A reliable deepfake detector or spoofing countermeasure (CM) should be robust in the face of unpredictable spoofing attacks. To encourage the learning of more generaliseable artefacts, rather than those specific only to known attacks, CMs are usually exposed to a broad variety of different attacks during training. Even so, the performance of deeplearning-based CM solutions are known to vary, sometimes substantially, when they are retrained with different initialisations, hyper-parameters or training data partitions. We show in this paper that the potency of spoofing attacks, also deep-learning-based, can similarly vary according to training conditions, sometimes resulting in substantial degradations to detection performance. Nevertheless, while a RawNet2 CM model is vulnerable when only modest adjustments are made to the attack algorithm, those based upon graph attention networks and self-supervised learning are reassuringly robust. The focus upon training data generated with different attack algorithms might not be sufficient on its own to ensure generaliability; some form of spoofing attack augmentation at the algorithm level can be complementary.
Wanying Ge, Xin Wang 0037, Junichi Yamagishi, Massimiliano Todisco, Nicholas W. D. Evans
ICASSP5
2024 Synvox2: Towards A Privacy-Friendly Voxceleb2 Dataset
abstract
The success of deep learning in speaker recognition relies heavily on the use of large datasets. However, the data-hungry nature of deep learning methods has already being questioned on account the ethical, privacy, and legal concerns that arise when using large-scale datasets of natural speech collected from real human speakers. For example, the widely-used VoxCeleb2 dataset for speaker recognition is no longer accessible from the official website. To mitigate these concerns, this work presents an initiative to generate a privacyfriendly synthetic VoxCeleb2 dataset that ensures the quality of the generated speech in terms of privacy, utility, and fairness. We also discuss the challenges of using synthetic data for the downstream task of speaker verification.
Xiaoxiao Miao, Xin Wang 0037, Erica Cooper, Junichi Yamagishi, Nicholas W. D. Evans, Massimiliano Todisco, Jean-François Bonastre, Mickael Rouvier
ICASSP5
2024 Speaker Anonymization Using Neural Audio Codec Language Models
abstract
The vast majority of approaches to speaker anonymization involve the extraction of fundamental frequency estimates, linguistic features and a speaker embedding which is perturbed to obfuscate the speaker identity before an anonymized speech waveform is resynthesized using a vocoder. Recent work has shown that x-vector transformations are difficult to control consistently: other sources of speaker information contained within fundamental frequency and linguistic features are re-entangled upon vocoding, meaning that anonymized speech signals still contain speaker information. We propose an approach based upon neural audio codecs (NACs), which are known to generate high-quality synthetic speech when combined with language models. NACs use quantized codes, which are known to effectively bottleneck speaker-related information: we demonstrate the potential of speaker anonymization systems based on NAC language modeling by applying the evaluation framework of the Voice Privacy Challenge 2022.
Michele Panariello, Francesco Nespoli, Massimiliano Todisco, Nicholas W. D. Evans
ICASSP4
2024 To what extent can ASV systems naturally defend against spoofing attacks?
Jee-Weon Jung, Xin Wang 0037, Nicholas W. D. Evans, Shinji Watanabe 0001, Hye-Jin Shim, Hemlata Tak, Siddhant Arora, Junichi Yamagishi, Joon Son Chung
INTERSPEECH3
2024 Harder or Different? Understanding Generalization of Audio Deepfake Detection
abstract
2705
Nicolas M. Müller, Nicholas W. D. Evans, Hemlata Tak, Philip Sperl, Konstantin Böttinger
INTERSPEECH2
2024 Spoof Diarization: "What Spoofed When" in Partially Spoofed Audio
Lin Zhang 0054, Xin Wang 0037, Erica Cooper, Mireia Díez, Federico Landini, Nicholas W. D. Evans, Junichi Yamagishi
INTERSPEECH6
2024 t-EER: Parameter-Free Tandem Evaluation of Countermeasures and Biometric Comparators
abstract
Presentation attack (spoofing) detection (PAD) typically operates alongside biometric verification to improve reliablity in the face of spoofing attacks. Even though the two sub-systems operate in tandem to solve the single task of reliable biometric verification, they address different detection tasks and are hence typically evaluated separately. Evidence shows that this approach is suboptimal. We introduce a new metric for the joint evaluation of PAD solutions operating in situ with biometric verification. In contrast to the tandem detection cost function proposed recently, the new tandem equal error rate (t-EER) is parameter free. The combination of two classifiers nonetheless leads to a set of operating points at which false alarm and miss rates are equal and also dependent upon the prevalence of attacks. We therefore introduce the concurrent t-EER, a unique operating point which is invariable to the prevalence of attacks. Using both modality (and even application) agnostic simulated scores, as well as real scores for a voice biometrics application, we demonstrate application of the t-EER to a wide range of biometric system evaluations under attack. The proposed approach is a strong candidate metric for the tandem evaluation of PAD systems and biometric comparators.
Tomi Kinnunen, Kong-Aik Lee, Hemlata Tak, Nicholas W. D. Evans, Andreas Nautsch
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 A False Sense of Privacy: Towards a Reliable Evaluation Methodology for the Anonymization of Biometric Data
abstract
Biometric data contains distinctive human traits such as facial features or gait patterns. The use of biometric data permits an individuation so exact that the data is utilized effectively in identification and authentication systems. But for this same reason, privacy protections become indispensably necessary. Privacy protection is extensively afforded by the technique of anonymization. Anonymization techniques protect sensitive personal data from biometrics by obfuscating or removing information that allows linking records to the generating individuals, to achieve high levels of anonymity. However, our understanding and possibility to develop effective anonymization relies, in equal parts, on the effectiveness of the methods employed to evaluate anonymization performance. In this paper, we assess the state-of-the-art methods used to evaluate the performance of anonymization techniques for facial images and for gait patterns. We demonstrate that the state-of-the-art evaluation methods have serious and frequent shortcomings. In particular, we find that the underlying assumptions of the state-of-the-art are quite unwarranted. State-of-the-art methods generally assume a difficult recognition scenario and thus a weak adversary. However, that assumption causes state-of-the-art evaluations to grossly overestimate the performance of the anonymization. Therefore, we propose a strong adversary which is aware of the anonymization in place. This adversary model implements an appropriate measure of anonymization performance. We improve the selection process for the evaluation dataset, and we reduce the numbers of identities contained in the dataset while ensuring that these identities remain easily distinguishable from one another. Our novel evaluation methodology surpasses the state-of-the-art because we measure worst-case performance and so deliver a highly reliable evaluation of biometric anonymization techniques.
Simon Hanisch, Julian Todt, Jose Patino 0001, Nicholas W. D. Evans, Thorsten Strufe
Proc. Priv. Enhancing Technol.4
2024 The VoicePrivacy 2022 Challenge: Progress and Perspectives in Voice Anonymisation
abstract
The VoicePrivacy Challenge promotes the development of voice anonymisation solutions for speech technology. In this paper we present a systematic overview and analysis of the second edition held in 2022. We describe the voice anonymisation task and datasets used for system development and evaluation, present the different attack models used for evaluation, and the associated objective and subjective metrics. We describe three anonymisation baselines, provide a summary description of the anonymisation systems developed by challenge participants, and report objective and subjective evaluation results for all. In addition, we describe post-evaluation analyses and a summary of related work reported in the open literature. Results show that solutions based on voice conversion better preserve utility, that an alternative which combines automatic speech recognition with synthesis achieves greater privacy, and that a privacy-utility trade-off remains inherent to current anonymisation solutions. Finally, we present our ideas and priorities for future VoicePrivacy Challenge editions.
Michele Panariello, Natalia A. Tomashenko, Xin Wang 0037, Xiaoxiao Miao, Pierre Champion, Hubert Nourtel, Massimiliano Todisco, Nicholas W. D. Evans, Emmanuel Vincent 0001, Junichi Yamagishi
IEEE ACM Trans. Audio Speech Lang. Process.8
2023 Can Spoofing Countermeasure And Speaker Verification Systems Be Jointly Optimised?
abstract
Spoofing countermeasure (CM) and automatic speaker verification (ASV) sub-systems can be used in tandem with a backend classifier as a solution to the spoofing aware speaker verification (SASV) task. The two sub-systems are typically trained independently to solve different tasks. While our previous work demonstrated the potential of joint optimisation, it also showed a tendency to over-fit to speakers and a lack of sub-system complementarity. Using only a modest quantity of auxiliary data collected from new speakers, we show that joint optimisation degrades the performance of separate CM and ASV sub-systems, but that it nonetheless improves complementarity, thereby delivering superior SASV performance. Using standard SASV evaluation data and protocols, joint optimisation reduces the equal error rate by 27% relative to performance obtained using fixed, independently-optimised subsystems under like-for-like training conditions.
Wanying Ge, Hemlata Tak, Massimiliano Todisco, Nicholas W. D. Evans
ICASSP4
2023 Towards Single Integrated Spoofing-aware Speaker Verification Embeddings
Sung Hwan Mun, Hye-Jin Shim, Hemlata Tak, Xin Wang 0037, Xuechen Liu 0001, Md. Sahidullah, Myeonghun Jeong, Min Hyun Han, Massimiliano Todisco, Kong-Aik Lee, Junichi Yamagishi, Nicholas W. D. Evans, Tomi Kinnunen, Nam Soo Kim, Jee-Weon Jung
INTERSPEECH12
2023 Malafide: a novel adversarial convolutive noise attack against deepfake and spoofing detection systems
Michele Panariello, Wanying Ge, Hemlata Tak, Massimiliano Todisco, Nicholas W. D. Evans
INTERSPEECH5
2023 Vocoder drift in x-vector-based speaker anonymization
Michele Panariello, Massimiliano Todisco, Nicholas W. D. Evans
INTERSPEECH3
2023 Range-Based Equal Error Rate for Spoof Localization
Lin Zhang 0054, Xin Wang 0037, Erica Cooper, Nicholas W. D. Evans, Junichi Yamagishi
INTERSPEECH4
2023 StressID: a Multimodal Dataset for Stress Identification
abstract
StressID is a new dataset specifically designed for stress identification fromunimodal and multimodal data. It contains videos of facial expressions, audiorecordings, and physiological signals. The video and audio recordings are acquiredusing an RGB camera with an integrated microphone. The physiological datais composed of electrocardiography (ECG), electrodermal activity (EDA), andrespiration signals that are recorded and monitored using a wearable device. Thisexperimental setup ensures a synchronized and high-quality multimodal data col-lection. Different stress-inducing stimuli, such as emotional video clips, cognitivetasks including mathematical or comprehension exercises, and public speakingscenarios, are designed to trigger a diverse range of emotional responses. Thefinal dataset consists of recordings from 65 participants who performed 11 tasks,as well as their ratings of perceived relaxation, stress, arousal, and valence levels.StressID is one of the largest datasets for stress identification that features threedifferent sources of data and varied classes of stimuli, representing more than39 hours of annotated data in total. StressID offers baseline models for stressclassification including a cleaning, feature extraction, and classification phase foreach modality. Additionally, we provide multimodal predictive models combiningvideo, audio, and physiological inputs. The data and the code for the baselines areavailable at https://project.inria.fr/stressid/.
Hava Chaptoukaev, Valeriya Strizhkova, Michele Panariello, Bianca Dalpaos, Aglind Reka, Valeria Manera, Susanne Thümmler, Esma Ismailova, Nicholas W. D. Evans, François Brémond, Massimiliano Todisco, Maria A. Zuluaga, Laura M. Ferrari
NeurIPS9
2023 ASVspoof 2021: Towards Spoofed and Deepfake Speech Detection in the Wild
abstract
Benchmarking initiatives support the meaningful comparison of competing solutions to prominent problems in speech and language processing. Successive benchmarking evaluations typically reflect a progressive evolution from ideal lab conditions towards to those encountered in the wild. ASVspoof, the spoofing and deepfake detection initiative and challenge series, has followed the same trend. This article provides a summary of the ASVspoof 2021 challenge and the results of 54 participating teams that submitted to the evaluation phase. For the logical access (LA) task, results indicate that countermeasures are robust to newly introduced encoding and transmission effects. Results for the physical access (PA) task indicate the potential to detect replay attacks in real, as opposed to simulated physical spaces, but a lack of robustness to variations between simulated and real acoustic environments. The Deepfake (DF) task, new to the 2021 edition, targets solutions to the detection of manipulated, compressed speech data posted online. While detection solutions offer some resilience to compression effects, they lack generalization across different source datasets. In addition to a summary of the top-performing systems for each task, new analyses of influential data factors and results for hidden data subsets, the article includes a review of post-challenge results, an outline of the principal challenge limitations and a road-map for the future of ASVspoof.
Xuechen Liu 0001, Xin Wang 0037, Md. Sahidullah, Jose Patino 0001, Héctor Delgado, Tomi Kinnunen, Massimiliano Todisco, Junichi Yamagishi, Nicholas W. D. Evans, Andreas Nautsch, Kong-Aik Lee
IEEE ACM Trans. Audio Speech Lang. Process.9
2023 The PartialSpoof Database and Countermeasures for the Detection of Short Fake Speech Segments Embedded in an Utterance
abstract
Automatic speaker verification is susceptible to various manipulations and spoofing, such as text-to-speech synthesis, voice conversion, replay, tampering, adversarial attacks, and so on. We consider a new spoofing scenario called "Partial Spoof" (PS) in which synthesized or transformed speech segments are embedded into a bona fide utterance. While existing countermeasures (CMs) can detect fully spoofed utterances, there is a need for their adaptation or extension to the PS scenario. We propose various improvements to construct a significantly more accurate CM that can detect and locate short-generated spoofed speech segments at finer temporal resolutions. First, we introduce newly developed self-supervised pre-trained models as enhanced feature extractors. Second, we extend our PartialSpoof database by adding segment labels for various temporal resolutions. Since the short spoofed speech segments to be embedded by attackers are of variable length, six different temporal resolutions are considered, ranging from as short as 20 ms to as large as 640 ms. Third, we propose a new CM that enables the simultaneous use of the segment-level labels at different temporal resolutions as well as utterance-level labels to execute utterance- and segment-level detection at the same time. We also show that the proposed CM is capable of detecting spoofing at the utterance level with low error rates in the PS scenario as well as in a related logical access (LA) scenario. The equal error rates of utterance-level detection on the PartialSpoof database and ASVspoof 2019 LA database were 0.77 and 0.90%, respectively.
Lin Zhang 0054, Xin Wang 0037, Erica Cooper, Nicholas W. D. Evans, Junichi Yamagishi
IEEE ACM Trans. Audio Speech Lang. Process.4
2022 Explaining Deep Learning Models for Spoofing and Deepfake Detection with Shapley Additive Explanations
abstract
International audience
Wanying Ge, Jose Patino 0001, Massimiliano Todisco, Nicholas W. D. Evans
ICASSP4
2022 AASIST: Audio Anti-Spoofing Using Integrated Spectro-Temporal Graph Attention Networks
abstract
Artefacts that differentiate spoofed from bona-fide utterances can reside in specific temporal or spectral intervals. Their reliable detection usually depends upon computationally demanding ensemble systems where each subsystem is tuned to some specific artefacts. We seek to develop an efficient, single system that can detect a broad range of different spoofing attacks without score-level ensembles. We propose a novel heterogeneous stacking graph attention layer that models artefacts spanning heterogeneous temporal and spectral intervals with a heterogeneous attention mechanism and a stack node. With a new max graph operation that involves a competitive mechanism and a new readout scheme, our approach, named AASIST, outperforms the current state-of-the-art by 20% relative. Even a lightweight variant, AASIST-L, with only 85k parameters, outperforms all competing systems.
Jee-Weon Jung, Hee-Soo Heo, Hemlata Tak, Hye-Jin Shim, Joon Son Chung, Bong-Jin Lee, Ha-Jin Yu, Nicholas W. D. Evans
ICASSP8
2022 Rawboost: A Raw Data Boosting and Augmentation Method Applied to Automatic Speaker Verification Anti-Spoofing
abstract
International audience
Hemlata Tak, Madhu R. Kamble, Jose Patino 0001, Massimiliano Todisco, Nicholas W. D. Evans
ICASSP5
2022 SASV 2022: The First Spoofing-Aware Speaker Verification Challenge
abstract
The first spoofing-aware speaker verification (SASV) challenge aims to integrate research efforts in speaker verification and anti-spoofing.We extend the speaker verification scenario by introducing spoofed trials to the usual set of target and impostor trials.In contrast to the established ASVspoof challenge where the focus is upon separate, independently optimised spoofing detection and speaker verification sub-systems, SASV targets the development of integrated and jointly optimised solutions.Pre-trained spoofing detection and speaker verification models are provided as open source and are used in two baseline SASV solutions.Both models and baselines are freely available to participants and can be used to develop back-end fusion approaches or end-to-end solutions.Using the provided common evaluation protocol, 23 teams submitted SASV solutions.When assessed with target, bona fide non-target and spoofed non-target trials, the top-performing system reduces the equal error rate of a conventional speaker verification system from 23.83% to 0.13%.SASV challenge results are a testament to the reliability of today's state-of-the-art approaches to spoofing detection and speaker verification.
Jee-Weon Jung, Hemlata Tak, Hye-Jin Shim, Hee-Soo Heo, Bong-Jin Lee, Soo-Whan Chung, Ha-Jin Yu, Nicholas W. D. Evans, Tomi Kinnunen
INTERSPEECH8
2022 Towards a unified assessment framework of speech pseudonymisation
abstract
Anonymisation and pseudonymisation are two similar concepts used in privacy preservation for speech data. With no established definitions for these tasks, nor standard approaches to assessment, this paper provides definitions and presents two complementary assessment frameworks. The first is based on voice similarity matrices which provide both an immediate visualisation of privacy protection performance at the speaker level and two objective measures in the form of de-identification and voice distinctiveness preservation. The approach readily highlights imbalances in system performance at the speaker level. The second, referred to as the zero evidence biometric recognition assessment (ZEBRA) framework, is based on information theory and measures the amount of private information disclosed in speech data. The paper presents also an extension to the original ZEBRA framework. It aims to reflect the robustness of the privacy safeguard when a privacy adversary adapts to the protected speech. We demonstrate the application of both frameworks to assess pseudonymisation performance on the two VoicePrivacy 2020 challenge baseline solutions plus a third one. The two frameworks were designed independently of each other. The ZEBRA framework is fully consistent with the Bayesian decision theory and the other framework focuses instead on speaker-wise visualisations of a system performance. Thus, while metrics derived from them bear similarities, they expose differences in safeguard behavior. The assessment of pseudonymisation remains challenging and merits greater attention in the future.
Paul-Gauthier Noé, Andreas Nautsch, Nicholas W. D. Evans, Jose Patino 0001, Jean-François Bonastre, Natalia A. Tomashenko, Driss Matrouf
Comput. Speech Lang.3
2022 The VoicePrivacy 2020 Challenge: Results and findings
Natalia A. Tomashenko, Xin Wang 0037, Emmanuel Vincent 0001, Jose Patino 0001, Brij Mohan Lal Srivastava, Paul-Gauthier Noé, Andreas Nautsch, Nicholas W. D. Evans, Junichi Yamagishi, Benjamin O'Brien, Anaïs Chanclu, Jean-François Bonastre, Massimiliano Todisco, Mohamed Maouche
Comput. Speech Lang.8
2021 Speaker Embeddings for Diarization of Broadcast Data In The Allies Challenge
abstract
Diarization consists in the segmentation of speech signals and the clustering of homogeneous speaker segments. State-of-the-art systems typically operate upon speaker embeddings, such as i-vectors or neural x-vectors, extracted from mel cepstral coefficients (MFCCs) or spectrograms. The recent SincNet architecture extracts x-vectors directly from raw speech signals. The work reported in this paper compares the performance of different embeddings extracted from MFCCs or the raw signal for speaker diarization and broadcast media treated with compression and sub-sampling, operations which typically degrade performance. Experiments are performed with the new ALLIES database that was designed to complement existing, publicly available French corpora of broadcast radio and TV shows. Results show that, in adverse conditions, with compression and sampling mismatch, SincNet x-vectors outperform i-vectors and x-vectors by relative DERs of 43% and 73% respectively. Additionally we found that SincNet x-vectors are not the absolute best embeddings but are more robust to data mismatch than others.
Anthony Larcher, Ambuj Mehrish, Marie Tahon, Sylvain Meignier, Jean Carrive, David Doukhan, Olivier Galibert, Nicholas W. D. Evans
ICASSP8
2021 End-to-End anti-spoofing with RawNet2
abstract
Spoofing countermeasures aim to protect automatic speaker verification systems from being manipulated by spoofed speech signals. While results from the most recent ASVspoof 2019 evaluation show great potential to detect most forms of attack, some continue to evade detection. This paper reports the first application of RawNet2 to anti-spoofing. RawNet2 ingests raw audio and has potential to learn cues that are not detectable using more traditional countermeasure solutions. We describe modifications made to the original RawNet2 architecture so that it can be applied to anti-spoofing. For A17 attacks, our RawNet2 systems results are the second-best reported, while the fusion of RawNet2 and baseline countermeasures gives the second-best results reported for the full ASVspoof 2019 logical access condition. Our results are reproducible with open source software.
Hemlata Tak, Jose Patino 0001, Massimiliano Todisco, Andreas Nautsch, Nicholas W. D. Evans, Anthony Larcher
ICASSP5
2021 Speaker Anonymisation Using the McAdams Coefficient
abstract
Anonymisation has the goal of manipulating speech signals in order to degrade the reliability of automatic approaches to speaker recognition, while preserving other aspects of speech, such as those relating to intelligibility and naturalness. This paper reports an approach to anonymisation that, unlike other current approaches, requires no training data, is based upon well-known signal processing techniques and is both efficient and effective. The proposed solution uses the McAdams coefficient to transform the spectral envelope of speech signals. Results derived using common VoicePrivacy 2020 databases and protocols show that random, optimised transformations can outperform competing solutions in terms of anonymisation while causing only modest, additional degradations to intelligibility, even in the case of a semi-informed privacy adversary.
Jose Patino 0001, Natalia A. Tomashenko, Massimiliano Todisco, Andreas Nautsch, Nicholas W. D. Evans
Interspeech5
2021 Privacy-Preserving Voice Anti-Spoofing Using Secure Multi-Party Computation
abstract
International audience
Oubaïda Chouchane, Baptiste Brossier, Jorge Esteban Gamboa Gamboa, Thomas Lardy, Hemlata Tak, Orhan Ermis, Madhu R. Kamble, Jose Patino 0001, Nicholas W. D. Evans, Melek Önen, Massimiliano Todisco
Interspeech9
2021 Partially-Connected Differentiable Architecture Search for Deepfake and Spoofing Detection
abstract
International audience
Wanying Ge, Michele Panariello, Jose Patino 0001, Massimiliano Todisco, Nicholas W. D. Evans
Interspeech5
2021 PANACEA Cough Sound-Based Diagnosis of COVID-19 for the DiCOVA 2021 Challenge
abstract
The COVID-19 pandemic has led to the saturation of public health services worldwide.In this scenario, the early diagnosis of SARS-Cov-2 infections can help to stop or slow the spread of the virus and to manage the demand upon health services.This is especially important when resources are also being stretched by heightened demand linked to other seasonal diseases, such as the flu.In this context, the organisers of the DiCOVA 2021 challenge have collected a database with the aim of diagnosing COVID-19 through the use of coughing audio samples.This work presents the details of the automatic system for COVID-19 detection from cough recordings presented by team PANACEA.This team consists of researchers from two European academic institutions and one company: EURECOM (France), University of Granada (Spain), and Biometric Vox S.L. (Spain).We developed several systems based on established signal processing and machine learning methods.Our best system employs a Teager energy operator cepstral coefficients (TECCs) based frontend and Light gradient boosting machine (LightGBM) backend.The AUC obtained by this system on the test set is 76.31% which corresponds to a 10% improvement over the official baseline.
Madhu R. Kamble, José A. González 0001, Teresa Grau, Juan M. Espín, Lorenzo Cascioli, Alejandro Gómez Alanís, Jose Patino 0001, Roberto Font, Antonio M. Peinado, Ángel M. Gómez, Nicholas W. D. Evans, Maria A. Zuluaga, Massimiliano Todisco
Interspeech12
2021 Visualizing Classifier Adjacency Relations: A Case Study in Speaker Verification and Voice Anti-Spoofing
abstract
Whether it be for results summarization, or the analysis of classifier fusion, some means to compare different classifiers can often provide illuminating insight into their behaviour, (dis)similarity or complementarity. We propose a simple method to derive 2D representation from detection scores produced by an arbitrary set of binary classifiers in response to a common dataset. Based upon rank correlations, our method facilitates a visual comparison of classifiers with arbitrary scores and with close relation to receiver operating characteristic (ROC) and detection error trade-off (DET) analyses. While the approach is fully versatile and can be applied to any detection task, we demonstrate the method using scores produced by automatic speaker verification and voice anti-spoofing systems. The former are produced by a Gaussian mixture model system trained with VoxCeleb data whereas the latter stem from submissions to the ASVspoof 2019 challenge.
Tomi Kinnunen, Andreas Nautsch, Md. Sahidullah, Nicholas W. D. Evans, Xin Wang 0037, Massimiliano Todisco, Héctor Delgado, Junichi Yamagishi, Kong-Aik Lee
Interspeech4
2021 Graph Attention Networks for Anti-Spoofing
abstract
The cues needed to detect spoofing attacks against automatic speaker verification are often located in specific spectral sub-bands or temporal segments. Previous works show the potential to learn these using either spectral or temporal self-attention mechanisms but not the relationships between neighbouring sub-bands or segments. This paper reports our use of graph attention networks (GATs) to model these relationships and to improve spoofing detection performance. GATs leverage a self-attention mechanism over graph structured data to model the data manifold and the relationships between nodes. Our graph is constructed from representations produced by a ResNet. Nodes in the graph represent information either in specific sub-bands or temporal segments. Experiments performed on the ASVspoof 2019 logical access database show that our GAT-based model with temporal attention outperforms all of our baseline single systems. Furthermore, GAT-based systems are complementary to a set of existing systems. The fusion of GAT-based models with more conventional countermeasures delivers a 47% relative improvement in performance compared to the best performing single GAT system.
Hemlata Tak, Jee-Weon Jung, Jose Patino 0001, Massimiliano Todisco, Nicholas W. D. Evans
Interspeech5
2021 An Initial Investigation for Detecting Partially Spoofed Audio
abstract
International audience
Lin Zhang 0054, Xin Wang 0037, Erica Cooper, Junichi Yamagishi, Jose Patino 0001, Nicholas W. D. Evans
Interspeech6
2020 Artificial Bandwidth Extension Using Conditional Variational Auto-encoders and Adversarial Learning
abstract
Artificial bandwidth extension (ABE) algorithms have been developed to estimate missing highband frequency components (4-8kHz) to improve quality of narrowband (0-4kHz) telephone calls. Most ABE solutions employ deep neural networks (DNNs) due to their well-known ability to model highly complex, non-linear relationship between narrowband and highband features. Generative models such as conditional variational auto-encoders (CVAEs) are capable of modelling complex data distributions via latent representation learning. This paper reports their application to ABE. CVAEs, form of directed, graphical models, are exploited to model the probability distribution of highband features conditioned on narrowband features. While CVAEs are trained with the standard mean square criterion (MSE), their combination with adversarial learning give further improvements. When compared to results obtained with the baseline approach, the wideband PESQ is improved significantly by 0.21 points. The performance is also compared on an automatic speech recognition (ASR) task on the TIMIT dataset where word error rate (WER) is decreased by an absolute value of 0.3%.
Pramod B. Bachhav, Massimiliano Todisco, Nicholas W. D. Evans
ICASSP3
2020 The Privacy ZEBRA: Zero Evidence Biometric Recognition Assessment
abstract
International audience
Andreas Nautsch, Jose Patino 0001, Natalia A. Tomashenko, Junichi Yamagishi, Paul-Gauthier Noé, Jean-François Bonastre, Massimiliano Todisco, Nicholas W. D. Evans
INTERSPEECH8
2020 Speech Pseudonymisation Assessment Using Voice Similarity Matrices
abstract
The proliferation of speech technologies and rising privacy legislation calls for the development of privacy preservation solutions for speech applications. These are essential since speech signals convey a wealth of rich, personal and potentially sensitive information. Anonymisation, the focus of the recent VoicePrivacy initiative, is one strategy to protect speaker identity information. Pseudonymisation solutions aim not only to mask the speaker identity and preserve the linguistic content, quality and naturalness, as is the goal of anonymisation, but also to preserve voice distinctiveness. Existing metrics for the assessment of anonymisation are ill-suited and those for the assessment of pseudonymisation are completely lacking. Based upon voice similarity matrices, this paper proposes the first intuitive visualisation of pseudonymisation performance for speech signals and two novel metrics for objective assessment. They reflect the two, key pseudonymisation requirements of de-identification and voice distinctiveness.
Paul-Gauthier Noé, Jean-François Bonastre, Driss Matrouf, Natalia A. Tomashenko, Andreas Nautsch, Nicholas W. D. Evans
INTERSPEECH6
2020 Spoofing Attack Detection Using the Non-Linear Fusion of Sub-Band Classifiers
abstract
International audience
Hemlata Tak, Jose Patino 0001, Andreas Nautsch, Nicholas W. D. Evans, Massimiliano Todisco
INTERSPEECH4
2020 Introducing the VoicePrivacy Initiative
abstract
The VoicePrivacy initiative aims to promote the development of privacy preservation tools for speech technology by gathering a new community to define the tasks of interest and the evaluation methodology, and benchmarking solutions through a series of challenges. In this paper, we formulate the voice anonymization task selected for the VoicePrivacy 2020 Challenge and describe the datasets used for system development and evaluation. We also present the attack models and the associated objective and subjective evaluation metrics. We introduce two anonymization baselines and report objective evaluation results.
Natalia A. Tomashenko, Brij Mohan Lal Srivastava, Xin Wang 0037, Emmanuel Vincent 0001, Andreas Nautsch, Junichi Yamagishi, Nicholas W. D. Evans, Jose Patino 0001, Jean-François Bonastre, Paul-Gauthier Noé, Massimiliano Todisco
INTERSPEECH7
2020 ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech
Xin Wang 0037, Junichi Yamagishi, Massimiliano Todisco, Héctor Delgado, Andreas Nautsch, Nicholas W. D. Evans, Md. Sahidullah, Ville Vestman, Tomi Kinnunen, Kong-Aik Lee, Lauri Juvela, Paavo Alku, Yu-Huai Peng, Hsin-Te Hwang, Yu Tsao 0001, Hsin-Min Wang, Sébastien Le Maguer, Zhen-Hua Ling
Comput. Speech Lang.6
2020 Tandem Assessment of Spoofing Countermeasures and Automatic Speaker Verification: Fundamentals
abstract
Recent years have seen growing efforts to develop spoofing countermeasures (CMs) to protect automatic speaker verification (ASV) systems from being deceived by manipulated or artificial inputs. The reliability of spoofing CMs is typically gauged using the equal error rate (EER) metric. The primitive EER fails to reflect application requirements and the impact of spoofing and CMs upon ASV and its use as a primary metric in traditional ASV research has long been abandoned in favour of risk-based approaches to assessment. This paper presents several new extensions to the tandem detection cost function (t-DCF), a recent risk-based approach to assess the reliability of spoofing CMs deployed in tandem with an ASV system. Extensions include a simplified version of the t-DCF with fewer parameters, an analysis of a special case for a fixed ASV system, simulations which give original insights into its interpretation and new analyses using the ASVspoof 2019 database. It is hoped that adoption of the t-DCF for the CM assessment will help to foster closer collaboration between the anti-spoofing and ASV research communities.
Tomi Kinnunen, Héctor Delgado, Nicholas W. D. Evans, Kong-Aik Lee, Ville Vestman, Andreas Nautsch, Massimiliano Todisco, Xin Wang 0037, Md. Sahidullah, Junichi Yamagishi, Douglas A. Reynolds
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 Latent Representation Learning for Artificial Bandwidth Extension Using a Conditional Variational Auto-encoder
abstract
Artificial bandwidth extension (ABE) algorithms can improve speech quality when wideband devices are used with narrowband devices or infrastructure. Most ABE solutions employ some form of memory, implying high-dimensional feature representations that increase both latency and complexity. Dimensionality reduction techniques have thus been developed to preserve efficiency. These entail the extraction of compact, low-dimensional representations that are then used with a standard regression model to estimate high-band components. Previous work shows that some form of supervision is crucial to the optimisation of dimensionality reduction techniques for ABE. This paper reports the first application of conditional variational auto-encoders (CVAEs) for supervised dimensionality reduction specifically tailored to ABE. CVAEs, form of directed, graphical models, are exploited to model higher-dimensional log-spectral data to extract the latent narrowband representations. When compared to results obtained with alternative dimensionality reduction techniques, objective and subjective assessments show that the probabilistic latent representations learned with CVAEs produce bandwidth-extended speech signals of notably better quality.
Pramod B. Bachhav, Massimiliano Todisco, Nicholas W. D. Evans
ICASSP3
2019 Privacy-Preserving Speaker Recognition with Cohort Score Normalisation
abstract
In many voice biometrics applications there is a requirement to preserve privacy, not least because of the recently enforced General Data Protection Regulation (GDPR). Though progress in bringing privacy preservation to voice biometrics is lagging behind developments in other biometrics communities, recent years have seen rapid progress, with secure computation mechanisms such as homomorphic encryption being applied successfully to speaker recognition. Even so, the computational overhead incurred by processing speech data in the encrypted domain is substantial. While still tolerable for single biometric comparisons, most state-of-the-art systems perform some form of cohort-based score normalisation, requiring many thousands of biometric comparisons. The computational overhead is then prohibitive, meaning that one must accept either degraded performance (no score normalisation) or potential for privacy violations. This paper proposes the first computationally feasible approach to privacy-preserving cohort score normalisation. Our solution is a cohort pruning scheme based on secure multi-party computation which enables privacy-preserving score normalisation using probabilistic linear discriminant analysis (PLDA) comparisons. The solution operates upon binary voice representations. While the binarisation is lossy in biometric rank-1 performance, it supports computationally-feasible biometric rank-n comparisons in the encrypted domain.
Andreas Nautsch, Jose Patino 0001, Amos Treiber, Themos Stafylakis, Petr Mizera, Massimiliano Todisco, Thomas Schneider 0003, Nicholas W. D. Evans
INTERSPEECH8
2019 The GDPR & Speech Data: Reflections of Legal and Technology Communities, First Steps Towards a Common Understanding
abstract
International audience
Andreas Nautsch, Catherine Jasserand, Els Kindt, Massimiliano Todisco, Isabel Trancoso, Nicholas W. D. Evans
INTERSPEECH6
2019 ASVspoof 2019: Future Horizons in Spoofed and Fake Audio Detection
abstract
ASVspoof, now in its third edition, is a series of community-led challenges which promote the development of countermeasures to protect automatic speaker verification (ASV) from the threat of spoofing. Advances in the 2019 edition include: (i) a consideration of both logical access (LA) and physical access (PA) scenarios and the three major forms of spoofing attack, namely synthetic, converted and replayed speech; (ii) spoofing attacks generated with state-of-the-art neural acoustic and waveform models; (iii) an improved, controlled simulation of replay attacks; (iv) use of the tandem detection cost function (t-DCF) that reflects the impact of both spoofing and countermeasures upon ASV reliability. Even if ASV remains the core focus, in retaining the equal error rate (EER) as a secondary metric, ASVspoof also embraces the growing importance of fake audio detection. ASVspoof 2019 attracted the participation of 63 research teams, with more than half of these reporting systems that improve upon the performance of two baseline spoofing countermeasures. This paper describes the 2019 database, protocols and challenge results. It also outlines major findings which demonstrate the real progress made in protecting against the threat of spoofing and fake audio.
Massimiliano Todisco, Xin Wang 0037, Ville Vestman, Md. Sahidullah, Héctor Delgado, Andreas Nautsch, Junichi Yamagishi, Nicholas W. D. Evans, Tomi Kinnunen, Kong-Aik Lee
INTERSPEECH8
2019 Preserving privacy in speaker and speech characterisation
abstract
Speech recordings are a rich source of personal, sensitive data that can be used to support a plethora of diverse applications, from health profiling to biometric recognition. It is therefore essential that speech recordings are adequately protected so that they cannot be misused. Such protection, in the form of privacy-preserving technologies, is required to ensure that: (i) the biometric profiles of a given individual (e.g., across different biometric service operators) are unlinkable; (ii) leaked, encrypted biometric information is irreversible, and that (iii) biometric references are renewable. Whereas many privacy-preserving technologies have been developed for other biometric characteristics, very few solutions have been proposed to protect privacy in the case of speech signals. Despite privacy preservation this is now being mandated by recent European and international data protection regulations. With the aim of fostering progress and collaboration between researchers in the speech, biometrics and applied cryptography communities, this survey article provides an introduction to the field, starting with a legal perspective on privacy preservation in the case of speech data. It then establishes the requirements for effective privacy preservation, reviews generic cryptography-based solutions, followed by specific techniques that are applicable to speaker characterisation (biometric applications) and speech characterisation (non-biometric applications). Glancing at non-biometrics, methods are presented to avoid function creep, preventing the exploitation of biometric information, e.g., to single out an identity in speech-assisted health care via speaker characterisation. In promoting harmonised research, the article also outlines common, empirical evaluation metrics for the assessment of privacy-preserving technologies for speech data.
Andreas Nautsch, Abelino Jiménez, Amos Treiber, Jascha Kolberg, Catherine Jasserand, Els Kindt, Héctor Delgado, Massimiliano Todisco, Mohamed Amine Hmani, Aymen Mtibaa, Mohammed Ahmed Abdelraheem, Alberto Abad, Francisco Teixeira, Driss Matrouf, Marta Gomez-Barrero, Dijana Petrovska-Delacrétaz, Gérard Chollet, Nicholas W. D. Evans, Christoph Busch 0001
Comput. Speech Lang.18
2018 Efficient Super-Wide Bandwidth Extension Using Linear Prediction Based Analysis-Synthesis
abstract
Many smart devices now support high-quality speech communication services at super-wide bandwidths. Often, however, speech quality is degraded when they are used with networks or devices which lack super-wideband support. Artificial bandwidth extension can then be used to improve speech quality. While approaches to wideband extension have been reported previously, this paper proposes an approach to super-wide bandwidth extension. The algorithm is based upon a classical source filter model in which spectral envelope and residual error information are extracted from a wideband signal using conventional linear prediction analysis. A form of spectral mirroring is then used to extend the residual error component before an extended super-wideband signal is derived from its combination with the original wideband envelope. Improvements to speech quality are confirmed with both objective and subjective assessments. These show that the quality of super-wideband speech, derived from the bandwidth extension of wideband speech, is comparable to that of speech processed with the standard enhanced voice services (EVS) codec with a bitrate of 13.2kbps. Without the need for statistical estimation of missing super-wideband components, the proposed algorithm is highly efficient and introduces only negligible latency.
Pramod B. Bachhav, Massimiliano Todisco, Nicholas W. D. Evans
ICASSP3
2018 Exploiting Explicit Memory Inclusion for Artificial Bandwidth Extension
abstract
Artificial bandwidth extension (ABE) algorithms have been developed to improve speech quality when wideband devices are used in conjunction with narrowband devices or infrastructure. While past work points to the benefit of using contextual information or memory for ABE, an understanding of the relative benefit of explicit memory inclusion, rather than just dynamic information, calls for a comparative, quantitative analysis. The need for practical ABE solutions calls further for the inclusion of memory without significant increases to latency or computational complexity. The paper reports the use of an information theoretic approach to show the potential of benefit of memory inclusion. Findings are validated through objective and subjective assessments of an ABE system which uses memory with only negligible increases to latency and computational complexity. Listening tests show that narrowband signals whose bandwidth is artificially extended with, rather than without the inclusion of memory, are of consistently improved quality.
Pramod B. Bachhav, Massimiliano Todisco, Nicholas W. D. Evans
ICASSP3
2018 Artificial Bandwidth Extension with Memory Inclusion Using Semi-supervised Stacked Auto-encoders
Pramod B. Bachhav, Massimiliano Todisco, Nicholas W. D. Evans
INTERSPEECH3
2018 Speech Database and Protocol Validation Using Waveform Entropy
Itshak Lapidot, Héctor Delgado, Massimiliano Todisco, Nicholas W. D. Evans, Jean-François Bonastre
INTERSPEECH4
2018 The EURECOM Submission to the First DIHARD Challenge
Jose Patino 0001, Héctor Delgado, Nicholas W. D. Evans
INTERSPEECH3
2018 Integrated Presentation Attack Detection and Automatic Speaker Verification: Common Features and Gaussian Back-end Fusion
abstract
International audience
Massimiliano Todisco, Héctor Delgado, Kong-Aik Lee, Md. Sahidullah, Nicholas W. D. Evans, Tomi Kinnunen, Junichi Yamagishi
INTERSPEECH5
2017 Artificial bandwidth extension using the constant Q transform
abstract
Most artificial bandwidth extension (ABE) algorithms are based on the classical source-filter model of speech production. This approach generally requires the dual extension of each component through independent processing. Alternative approaches reported recently operate on the spectrum. With human perception thought to be largely insensitive to phase, most such approaches focus on the extension of the magnitude spectrum alone and rely on Fourier spectral analysis. This paper reports an approach to ABE based on the constant Q transform (CQT), a more perceptually motivated approach to spectral analysis. A Gaussian mixture model is used to estimate missing highband components from available narrowband components before resynthesis with phase estimates obtained from the upsampled narrowband signal. Objective assessment shows that energy normalisation is critical to performance. These findings and the appeal of CQT for ABE are confirmed through informal subjective tests based on the mean opinion score.
Pramod B. Bachhav, Massimiliano Todisco, Moctar Mossi Idrissa, Christophe Beaugeant, Nicholas W. D. Evans
ICASSP5
2017 RedDots replayed: A new replay spoofing attack corpus for text-dependent speaker verification research
abstract
This paper describes a new database for the assessment of automatic speaker verification (ASV) vulnerabilities to spoofing attacks. In contrast to other recent data collection efforts, the new database has been designed to support the development of replay spoofing countermeasures tailored towards the protection of text-dependent ASV systems from replay attacks in the face of variable recording and playback conditions. Derived from the re-recording of the original RedDots database, the effort is aligned with that in text-dependent ASV and thus well positioned for future assessments of replay spoofing countermeasures, not just in isolation, but in integration with ASV. The paper describes the database design and re-recording, a protocol and some early spoofing detection results. The new “RedDots Replayed” database is publicly available through a creative commons license.
Tomi Kinnunen, Md. Sahidullah, Mauro Falcone, Luca Costantini, Rosa González Hautamäki, Dennis Alexander Lehmann Thomsen, Achintya Kumar Sarkar, Zheng-Hua Tan, Héctor Delgado, Massimiliano Todisco, Nicholas W. D. Evans, Ville Hautamäki, Kong-Aik Lee
ICASSP11
2017 The ASVspoof 2017 Challenge: Assessing the Limits of Replay Spoofing Attack Detection
abstract
The ASVspoof initiative was created to promote the development of countermeasures which aim to protect automatic speaker verification (ASV) from spoofing attacks. The first community-led, common evaluation held in 2015 focused on countermeasures for speech synthesis and voice conversion spoofing attacks. Arguably, however, it is replay attacks which pose the greatest threat. Such attacks involve the replay of recordings collected from enrolled speakers in order to provoke false alarms and can be mounted with greater ease using everyday consumer devices. ASVspoof 2017, the second in the series, hence focused on the development of replay attack countermeasures. This paper describes the database, protocols and initial findings. The evaluation entailed highly heterogeneous acoustic recording and replay conditions which increased the equal error rate (EER) of a baseline ASV system from 1.76% to 30.71%. Submissions were received from 49 research teams, 20 of which improved upon a baseline replay spoofing detector EER of 24.65%, in terms of replay/non-replay discrimination. While largely successful, the evaluation indicates that the quest for countermeasures which are resilient in the face of variable replay attacks remains very much alive.
Tomi Kinnunen, Md. Sahidullah, Héctor Delgado, Massimiliano Todisco, Nicholas W. D. Evans, Junichi Yamagishi, Kong-Aik Lee
INTERSPEECH5
2017 The I4U Mega Fusion and Collaboration for NIST Speaker Recognition Evaluation 2016
abstract
18th Annual Conference of the International Speech Communication Association, INTERSPEECH 2017, Stockholm, Sweden, 20-24 August 2017
Kong-Aik Lee, Ville Hautamäki, Tomi Kinnunen, Anthony Larcher, Andreas Nautsch, Themos Stafylakis, Gang Liu 0001, Mickael Rouvier, Wei Rao 0002, Federico Alegre, Man-Wai Mak, Achintya Kumar Sarkar, Héctor Delgado, Rahim Saeidi, Hagai Aronowitz, Aleksandr Sizov, Hanwu Sun, Trung Hieu Nguyen 0001, Guangsen Wang, Bin Ma 0001, Ville Vestman, Md. Sahidullah, M. Halonen, Anssi Kanervisto, Gaël Le Lan, Fahimeh Bahmaninezhad, Sergey Isadskiy, Christian Rathgeb, Christoph Busch 0001, Georgios Tzimiropoulos, Q. Qian, Q. Zhao, J. Xue, R. Jin, T. Zhao, Pierre-Michel Bousquet, Moez Ajili, Waad Ben Kheder, Driss Matrouf, Zhi Hao Lim, Chenglin Xu, Haihua Xu 0001, Chng Eng Siong, Benoit G. B. Fauve, Kaavya Sriskandaraja, Vidhyasaharan Sethu, W. W. Lin, Dennis Alexander Lehmann Thomsen, Zheng-Hua Tan, Massimiliano Todisco, Nicholas W. D. Evans, Haizhou Li 0001, John H. L. Hansen, Jean-François Bonastre, Eliathamby Ambikairajah
INTERSPEECH58
2017 Constant Q cepstral coefficients: A spoofing countermeasure for automatic speaker verification
Massimiliano Todisco, Héctor Delgado, Nicholas W. D. Evans
Comput. Speech Lang.3
2016 Utterance Verification for Text-Dependent Speaker Recognition: A Comparative Assessment Using the RedDots Corpus
abstract
Text-dependent automatic speaker verification naturally calls for the simultaneous verification of speaker identity and spoken content. These two tasks can be achieved with automatic speaker verification (ASV) and utterance verification (UV) technologies. While both have been addressed previously in the literature, a treatment of simultaneous speaker and utterance verification with a modern, standard database is so far lacking. This is despite the burgeoning demand for voice biometrics in a plethora of practical security applications. With the goal of improving overall verification performance, this paper reports different strategies for simultaneous ASV and UV in the context of short-duration, text-dependent speaker verification. Experiments performed on the recently released RedDots corpus are reported for three different ASV systems and four different UV systems. Results show that the combination of utterance verification with automatic speaker verification is (almost) universally beneficial with significant performance improvements being observed.
Tomi Kinnunen, Md. Sahidullah, Ivan Kukanov, Héctor Delgado, Massimiliano Todisco, Achintya Kumar Sarkar, Nicolai Bæk Thomsen, Ville Hautamäki, Nicholas W. D. Evans, Zheng-Hua Tan
INTERSPEECH9
2016 Integrated Spoofing Countermeasures and Automatic Speaker Verification: An Evaluation on ASVspoof 2015
abstract
It is well known that automatic speaker verification (ASV) systems can be vulnerable to spoofing. The community has responded to the threat by developing dedicated countermeasures aimed at detecting spoofing attacks. Progress in this area has accelerated over recent years, partly as a result of the first standard evaluation, ASVspoof 2015, which focused on spoofing detection in isolation from ASV. This paper investigates the integration of state-of-the-art spoofing countermeasures in combination with ASV. Two general strategies to countermeasure integration are reported: cascaded and parallel. The paper reports the first comparative evaluation of each approach performed with the ASVspoof 2015 corpus. Results indicate that, even in the case of varying spoofing attack algorithms, ASV performance remains robust when protected with a diverse set of integrated countermeasures.
Md. Sahidullah, Héctor Delgado, Massimiliano Todisco, Hong Yu 0015, Tomi Kinnunen, Nicholas W. D. Evans, Zheng-Hua Tan
INTERSPEECH6
2016 Articulation Rate Filtering of CQCC Features for Automatic Speaker Verification
Massimiliano Todisco, Héctor Delgado, Nicholas W. D. Evans
INTERSPEECH3
2016 On the Influence of Text Content on Pass-Phrase Strength for Short-Duration Text-Dependent Automatic Speaker Authentication
Giacomo Valenti, Adrien Daniel, Nicholas W. D. Evans
INTERSPEECH3
2016 Further optimisations of constant Q cepstral processing for integrated utterance and text-dependent speaker verification
abstract
Many authentication applications involving automatic speaker verification (ASV) demand robust performance using short-duration, fixed or prompted text utterances. Text constraints not only reduce the phone-mismatch between enrolment and test utterances, which generally leads to improved performance, but also provide an ancillary level of security. This can take the form of explicit utterance verification (UV). An integrated UV + ASV system should then verify access attempts which contain not just the expected speaker, but also the expected text content. This paper presents such a system and introduces new features which are used for both UV and ASV tasks. Based upon multi-resolution, spectro-temporal analysis and when fused with more traditional parameterisations, the new features not only generally outperform Mel-frequency cepstral coefficients, but also are shown to be complementary when fusing systems at score level. Finally, the joint operation of UV and ASV greatly decreases false acceptances for unmatched text trials.
Héctor Delgado, Massimiliano Todisco, Md. Sahidullah, Achintya Kumar Sarkar, Nicholas W. D. Evans, Tomi Kinnunen, Zheng-Hua Tan
SLT5
2016 An assessment of automatic speaker verification vulnerabilities to replay spoofing attacks
abstract
Abstract This paper analyses the threat of replay spoofing or presentation attacks in the context of automatic speaker verification. As relatively high‐technology attacks, speech synthesis and voice conversion, which have thus far received far greater attention in the literature, are probably beyond the means of the average fraudster. The implementation of replay attacks, in contrast, requires no specific expertise nor sophisticated equipment. Replay attacks are thus likely to be the most prolific in practice, while their impact is relatively under‐researched. The work presented here aims to compare at a high level the threat of replay attacks with those of speech synthesis and voice conversion. The comparison is performed using strictly controlled protocols and with six different automatic speaker verification systems including a state‐of‐the‐art iVector/probabilistic linear discriminant analysis system. Experiments show that low‐effort replay attacks present at least a comparable threat to speech synthesis and voice conversion. The paper also describes and assesses two replay attack countermeasures. A relatively new approach based on the local binary pattern analysis of speech spectrograms is shown to outperform a competing approach based on the detection of far‐field recordings. Copyright © 2016 John Wiley & Sons, Ltd.
Artur Janicki, Federico Alegre, Nicholas W. D. Evans
Secur. Commun. Networks3
2015 Non-linear acoustic echo cancellation using empirical mode decomposition
abstract
The increasing popularity of miniature devices and loudspeakers has fuelled research in non-linear acoustic echo cancellation (NAEC). This paper reports a novel approach to NAEC based on empirical mode decomposition (EMD), a recently developed technique in non-linear and non-stationary signal analysis. EMD decomposes any signal into a finite number of time varying sub-band signals termed intrinsic mode functions (IMFs). The new approach to NAEC presented here incorporates this multi-resolution analysis with conventional power filtering to estimate non-linear echo in each IMF. Comparative experiments with a competitive baseline approach to NAEC based on pure power filtering show that the new EMD approach achieves greater non-linear echo reduction and faster convergence.
Leela K. Gudupudi, Navin Chatlani, Christophe Beaugeant, Nicholas W. D. Evans
ICASSP4
2015 ASVspoof 2015: the first automatic speaker verification spoofing and countermeasures challenge
abstract
An increasing number of independent studies have con-firmed the vulnerability of automatic speaker verification (ASV) technology to spoofing. However, in comparison to that involving other biometric modalities, spoofing and countermea-sure research for ASV is still in its infancy. A current barrier to progress is the lack of standards which impedes the comparison of results generated by different researchers. The ASVspoof ini-tiative aims to overcome this bottleneck through the provision of standard corpora, protocols and metrics to support a common evaluation. This paper introduces the first edition, summaries the results and discusses directions for future challenges and re-search.
Zhizheng Wu 0001, Tomi Kinnunen, Nicholas W. D. Evans, Junichi Yamagishi, Cemal Hanilçi, Md. Sahidullah, Aleksandr Sizov
INTERSPEECH3
2015 Automatic speaker verification spoofing and countermeasures (ASVspoof 2015): open discussion and future plans
Junichi Yamagishi, Nicholas W. D. Evans
INTERSPEECH2
2015 Spoofing and countermeasures for speaker verification: A survey
Zhizheng Wu 0001, Nicholas W. D. Evans, Tomi Kinnunen, Junichi Yamagishi, Federico Alegre, Haizhou Li 0001
Speech Commun.2
2015 Guest Editorial Special Issue on Biometric Spoofing and Countermeasures
abstract
While biometrics technology has created new solutions to person authentication and has evolved to play a critical role in personal, national, and global security, the potential for the technology to be fooled orspoofedis now widely acknowledged. For example, fingerprint verification systems can be spoofed with a synthetic material, such as gelatine, inscribed with the fingerprint ridges of an enrolled individual. Iris and face recognition systems are vulnerable to printed photographs or video sequences of an enrolled user’s eye or face. Speaker recognition systems can be spoofed through the use of replayed, synthesized, or converted speech.
Nicholas W. D. Evans, Stan Z. Li, Sébastien Marcel, Arun Ross
IEEE Trans. Inf. Forensics Secur.1
2014 Evasion and obfuscation in automatic speaker verification
abstract
The potential for biometric systems to be manipulated through some form of subversion is well acknowledged. One such approach known as spoofing relates to the provocation of false accepts in authentication applications. Another approach referred to as obfuscation relates to the provocation of missed detections in surveillance applications. While the automatic speaker verification research community is now addressing spoofing and countermeasures, vulnerabilities to obfuscation remain largely unknown. This paper reports the first study. Our work with standard NIST datasets and protocols shows that the equal error rate of a standard GMM-UBM system is increased from 9% to 48% through obfuscation, whereas that of a state-of-the-art i-vector system increases from 3% to 20%. We also present a generalised approach to obfuscation detection which succeeds in detecting almost all attempts to evade detection.
Federico Alegre, Giovanni Soldi, Nicholas W. D. Evans
ICASSP3
2014 A subspace co-training framework for multi-view clustering
Xuran Zhao, Nicholas W. D. Evans, Jean-Luc Dugelay
Pattern Recognit. Lett.2
2013 Spoofing countermeasures to protect automatic speaker verification from voice conversion
abstract
This paper presents a new countermeasure for the protection of automatic speaker verification systems from spoofed, converted voice signals. The new countermeasure exploits the common shift applied to the spectral slope of consecutive speech frames involved in the mapping of a spoofer's voice signal towards a statistical model of a given target. While the countermeasure exploits prior knowledge of the attack in an admittedly unrealistic sense, it is shown to detect almost all spoofed signals which otherwise provoke significant increases in false acceptance. The work also discusses the need for formal evaluations to develop new countermeasures which are less reliant on prior knowledge.
Federico Alegre, Asmaa Amehraye, Nicholas W. D. Evans
ICASSP3
2013 Voice activity detection based on a statistical semiparametric test
abstract
This paper adresses the voice activity detection problem within a semiparametric hypothesis testing framework. Semiparametric detection consists in combining the statistical optimality of a parametric test with the robustness regarding the learning data of a nonparametric test. The proposed semiparametric approach splits the frame vector into two parts such that the first part has a known statistical distribution. The second part is processed by a non-parametric detector producing a binary decision. A likelihood ratio test, based on the first part and the nonparametric binary decision, is then applied to classify the frame as either speech or nonspeech. The statistical performance of the resulting fusion test is analytically established and validated using real speech signals.
Asmaa Amehraye, Lionel Fillatre, Nicholas W. D. Evans
ICASSP3
2013 An experimental framework for the derivation of perceptually-optimal noise suppression functions
abstract
This paper presents a novel experimental framework designed to derive, through subjective testings, noise suppression functions which are perceptually optimal under specific experimental conditions. Noisy speech sequences are continuously processed according to a gain curve function of the a priori SNR that listeners are required to adjust two points at a time with respect to specified perceptual criteria. An experiment based on this framework is reported testing one specific combination of speech and noise signals. The specified perceptual criterion was the suitability for a phone conversation. The resulting mean experimental gain function shows a statistically significant deviation from an ideal Wiener filter. Experiments based on this framework are repeatable, suit untrained listeners and are considerably faster than conventional subjective testing methods, without the necessity to place restrictive assumptions on the assessed noise suppression function.
Adrien Daniel, Ludovick Lepauloux, Christelle Yemdji, Nicholas W. D. Evans, Christophe Beaugeant
ICASSP4
2013 Open-set semi-supervised audio-visual speaker recognition using co-training LDA and Sparse Representation Classifiers
abstract
Semi-supervised learning is attracting growing interest within the biometrics community. Almost all prior work focuses on closed-set scenarios, in which samples labelled automatically are assumed to belong to an enrolled class. This is often not the case in realistic applications and thus open-set alternatives are needed. This paper proposes a new approach to open-set, semi-supervised learning based on co-training, Linear Discriminant Analysis (LDA) subspaces and Sparse Representation Classifiers (SRCs). Experiments on the standard MOBIO dataset show how the new approach can utilize automatically labelled data to augment a smaller, manually labelled dataset and thus improve the performance of an open-set audio-visual person recognition system.
Xuran Zhao, Nicholas W. D. Evans, Jean-Luc Dugelay
ICASSP2
2013 A new speaker verification spoofing countermeasure based on local binary patterns
abstract
This paper presents a new countermeasure for the protection of automatic speaker verification systems from spoofed, converted voice signals.The new countermeasure is based on the analysis of a sequence of acoustic feature vectors using Local Binary Patterns (LBPs).Compared to existing approaches the new countermeasure is less reliant on prior knowledge and affords robust protection from not only voice conversion, for which it is optimised, but also spoofing attacks from speech synthesis and artificial signals, all of which otherwise provoke significant increases in false acceptance.The work highlights the difficulty in detecting converted voice and also discusses the need for formal evaluations to develop new countermeasures which are less reliant on prior knowledge and thus more reflective of practical use cases.
Federico Alegre, Ravichander Vipperla, Asmaa Amehraye, Nicholas W. D. Evans
INTERSPEECH4
2013 Spoofing and countermeasures for automatic speaker verification
abstract
It is widely acknowledged that most biometric systems are vulnerable to spoofing, also known as imposture.While vulnerabilities and countermeasures for other biometric modalities have been widely studied, e.g.face verification, speaker verification systems remain vulnerable.This paper describes some specific vulnerabilities studied in the literature and presents a brief survey of recent work to develop spoofing countermeasures.The paper concludes with a discussion on the need for standard datasets, metrics and formal evaluations which are needed to assess vulnerabilities to spoofing in realistic scenarios without prior knowledge.
Nicholas W. D. Evans, Tomi Kinnunen, Junichi Yamagishi
INTERSPEECH1
2013 Using linguistic information to detect overlapping speech
abstract
Overlapping speech is still a major cause of error in many speech processing applications, currently without any satisfactory solution. This paper considers the problem of detecting segments of overlapping speech within meeting recordings. Using an HMM-based framework recordings are segmented into intervals containing non-speech, speech and overlapping speech. New to this contribution is the use of linguistic information, where spoken content is used to improve overlap detection. Using language models for speech and overlap, an overlap score is created for every spoken word and used as an additional feature within the HMM framework. Experiments conducted on the AMI corpus demonstrate the potential of the proposed linguistic features.
Jürgen T. Geiger, Florian Eyben, Nicholas W. D. Evans, Björn W. Schuller, Gerhard Rigoll
INTERSPEECH3
2012 Dual amplifier and loudspeaker compensation using fast convergent and cascaded approaches to non-linear acoustic echo cancellation
abstract
This paper focuses on cascaded approaches to non-linear acoustic echo cancellation (AEC) for mobile communications. The contributions in this paper are two-fold. They relate (i) to computationally efficient pre-processing and clipping compensation which aims to improve non-linear modelling and (ii) decorrelation filtering which aims to improve the tracking performance of a conventional linear AEC algorithm. While well-established in the literature the two modules require significant development in order that they function coherently in a cascaded approach. This paper presents new, adaptive parameterisation procedures for both modules and demonstrates significant improvements in terms of echo return loss enhancement when the two modules are combined.
Moctar Mossi Idrissa, Christelle Yemdji, Nicholas W. D. Evans, Christophe Beaugeant, Fabrice Plante, Fatimazahra Marfouq
ICASSP3
2012 Speech overlap detection and attribution using convolutive non-negative sparse coding
abstract
Overlapping speech is known to degrade speaker diarization performance with impacts on speaker clustering and segmentation. While previous work made important advances in detecting overlapping speech intervals and in attributing them to relevant speakers, the problem remains largely unsolved. This paper reports the first application of convolutive non-negative sparse coding (CNSC) to the overlap problem. CNSC aims to decompose a composite signal into its underlying contributory parts and is thus naturally suited to overlap detection and attribution. Experimental results on NIST RT data show that the CNSC approach gives comparable results to a state-of-the-art hidden Markov model based overlap detector. In a practical diarization system, CNSC based speaker attribution is shown to reduce the speaker error by over 40% relative in overlapping segments.
Ravichander Vipperla, Jürgen T. Geiger, Simon Bozonnet, Dong Wang 0013, Nicholas W. D. Evans, Björn W. Schuller, Gerhard Rigoll
ICASSP5
2012 CO-LDA: A Semi-supervised Approach to Audio-Visual Person Recognition
abstract
Client models used in Automatic Speaker Recognition (ASR) and Automatic Face Recognition (AFR) are usually trained with labelled data acquired in a small number of menthol sessions. The amount of training data is rarely sufficient to reliably represent the variation which occurs later during testing. Larger quantities of client-specific training data can always be obtained, but manual collection and labelling is often cost-prohibitive. Co-training, a paradigm of semi-supervised machine learning, which can exploit unlabelled data to enhance weakly learned client models. In this paper, we propose a co-LDA algorithm which uses both labelled and unlabelled data to capture greater intersession variation and to learn discriminative subspaces in which test examples can be more accurately classified. The proposed algorithm is naturally suited to audio-visual person recognition because vocal and visual biometric features intrinsically satisfy the assumptions of feature sufficiency and independency which guarantee the effectiveness of co-training. When tested on the MOBIO database, the proposed co-training system raises a baseline identification rate from 71% to 99% while in a verification task the Equal Error Rate (EER) is reduced from 18% to about 1%. To our knowledge, this is the first successful application of co-training in audio-visual biometric systems.
Xuran Zhao, Nicholas W. D. Evans, Jean-Luc Dugelay
ICME2
2012 Spoofing countermeasures for the protection of automatic speaker recognition systems against attacks with artificial signals
abstract
The vulnerability of automatic speaker recognition systems to imposture or spoofing is widely acknowledged. This paper shows that extremely high false alarm rates can be provoked by simple spoofing attacks with artificial, non-speech-like signals and highlights the need for spoofing countermeasures. We show that two new, but trivial countermeasures based on higher-level, dynamic features and voice quality assessment offer varying degrees of protection and that further work is needed to develop more robust spoofing countermeasure mechanisms. Finally, we show that certain classifiers are inherently more robust to such attacks than others which strengthens the case for fused-system approaches to automatic speaker recognition.
Federico Alegre, Ravichander Vipperla, Nicholas W. D. Evans
INTERSPEECH3
2012 Phone Adaptive Training for Speaker Diarization
abstract
The linguistic content of a speech signal is a source of unwanted variation which can degrade speaker diarization performance.This paper presents our latest work to reduce its impact.The new approach, referred to as Phone Adaptive Training (PAT), is analogous to speaker adaptive training used in automatic speech recognition.We report an oracle experiment which shows that PAT has the potential to deliver a 33% relative improvement in the diarization error rate of our baseline system.Practical experiments show significant improvements across two standard, independent evaluation datasets.
Simon Bozonnet, Ravichander Vipperla, Nicholas W. D. Evans
INTERSPEECH3
2012 Convolutive Non-Negative Sparse Coding and New Features for Speech Overlap Handling in Speaker Diarization
abstract
The effective handling of overlapping speech is at the limits of the current state of the art in speaker diarization.This paper presents our latest work in overlap detection.We report the combination of features derived through convolutive nonnegative sparse coding and new energy, spectral and voicingrelated features within a conventional HMM system.Overlap detection results are fully integrated into our top-down diarization system through the application of overlap exclusion and overlap labeling.Experiments on a subset of the AMI corpus show that the new system delivers significant reductions in missed speech and speaker error.Through overlap exclusion and labelling the overall diarization error rate is shown to improve by 6.4 % relative.
Jürgen T. Geiger, Ravichander Vipperla, Simon Bozonnet, Nicholas W. D. Evans, Björn W. Schuller, Gerhard Rigoll
INTERSPEECH4
2012 A Comparative Study of Bottom-Up and Top-Down Approaches to Speaker Diarization
abstract
This paper presents a theoretical framework to analyze the relative merits of the two most general, dominant approaches to speaker diarization involving bottom-up and top-down hierarchical clustering. We present an original qualitative comparison which argues how the two approaches are likely to exhibit different behavior in speaker inventory optimization and model training: bottom-up approaches will capture comparatively purer models and will thus be more sensitive to nuisance variation such as that related to the speech content; top-down approaches, in contrast, will produce less discriminative speaker models but, importantly, models which are potentially better normalized against nuisance variation. We report experiments conducted on two standard, single-channel NIST RT evaluation datasets which validate our hypotheses. Results show that competitive performance can be achieved with both bottom-up and top-down approaches (average DERs of 21% and 22%), and that neither approach is superior. Speaker purification, which aims to improve speaker discrimination, gives more consistent improvements with the top-down system than with the bottom-up system (average DERs of 19% and 25%), thereby confirming that the top-down system is less discriminative and that the bottom-up system is less stable. Finally, we report a new combination strategy that exploits the merits of the two approaches. Combination delivers an average DER of 17% and confirms the intrinsic complementary of the two approaches.
Nicholas W. D. Evans, Simon Bozonnet, Dong Wang 0013, Corinne Fredouille, Raphaël Troncy
IEEE Trans. Speech Audio Process.1
2012 Speaker Diarization: A Review of Recent Research
abstract
Speaker diarization is the task of determining “who spoke when?” in an audio or video recording that contains an unknown amount of speech and also an unknown number of speakers. Initially, it was proposed as a research topic related to automatic speech recognition, where speaker diarization serves as an upstream processing step. Over recent years, however, speaker diarization has become an important key technology for many tasks, such as navigation, retrieval, or higher level inference on audio data. Accordingly, many important improvements in accuracy and robustness have been reported in journals and conferences in the area. The application domains, from broadcast news, to lectures and meetings, vary greatly and pose different problems, such as having access to multiple microphones and multimodal information or overlapping speech. The most recent review of existing technology dates back to 2006 and focuses on the broadcast news domain. In this paper, we review the current state-of-the-art, focusing on research developed since 2006 that relates predominantly to speaker diarization for conference meetings. Finally, we present an analysis of speaker diarization performance as reported through the NIST Rich Transcription evaluations on meeting data and identify important areas for future research.
Xavier Anguera Miró, Simon Bozonnet, Nicholas W. D. Evans, Corinne Fredouille, Gerald Friedland, Oriol Vinyals
IEEE Trans. Speech Audio Process.3
2012 Direct posterior confidence for out-of-vocabulary spoken term detection
abstract
Spoken term detection (STD) is a key technology for spoken information retrieval. As compared to the conventional speech transcription and keyword spotting, STD is an open-vocabulary task and has to address out-of-vocabulary (OOV) terms. Approaches based on subword units, for example phones, are widely used to solve the OOV issue; however, performance on OOV terms is still substantially inferior to that of in-vocabulary (INV) terms. The performance degradation on OOV terms can be attributed to a multitude of factors. One particular factor we address in this article is the unreliable confidence estimation caused by weak acoustic and language modeling due to the absence of OOV terms in the training corpora. We propose a direct posterior confidence derived from a discriminative model, such as multilayer perceptron (MLP). The new confidence considers a wide-range acoustic context which is usually important for speech recognition and retrieval; moreover, it localizes on detected speech segments and therefore avoids the impact of long-span word context which is usually unreliable for OOV term detection. In this article, we first develop an extensive discussion about the modeling weakness problem associated with OOV terms, and then propose our approach to address this problem based on direct poster confidence. Our experiments carried out on spontaneous and conversational multiparty meeting speech, demonstrate that the proposed technique provides a significant improvement in STD performance as compared to conventional lattice-based confidence, in particular for OOV terms. Furthermore, the new confidence estimation approach is fused with other advanced techniques for OOV treatment, such as stochastic pronunciation modeling and discriminative confidence normalization. This leads to an integrated solution for OOV term detection that results in a large performance improvement.
Dong Wang 0013, Simon King 0001, Joe Frankel, Ravichander Vipperla, Nicholas W. D. Evans, Raphaël Troncy
ACM Trans. Inf. Syst.5
2011 Linguistic influences on bottom-up and top-down clustering for speaker diarization
abstract
While bottom-up approaches have emerged as the standard, default approach to clustering for speaker diarization we have always found the top-down approach gives equivalent or superior performance. Our recent work shows that significant gains in performance can be obtained when cluster purification is applied to the output of top down systems but that it can degrade performance when applied to the output of bottom-up systems. This paper demonstrates that these observations can be accounted for by factors unrelated to the speaker and that they can impact more strongly on the performance of bottom-up clustering strategies than top-down strategies. Experimental results confirm that clusters produced through top-down clustering are better normalized against phone variation than those produced through bottom-up clustering and that this accounts for the observed inconsistencies in purification performance. The work highlights the need for marginalization strategies which should encourage convergence toward different speakers rather than toward nuisance factors such as that those related to the linguistic content.
Simon Bozonnet, Dong Wang 0013, Nicholas W. D. Evans, Raphaël Troncy
ICASSP3
2011 Robust and low-cost cascaded non-linear acoustic echo cancellation
abstract
This paper addresses die problem of acoustic echo cancellation in non-linear environments. The first contribution relates to tile use of a cascaded model which divides the loudspeaker enclosure microphone system into two main blocks; the first models the down link transducers which are assumed to be the main source of non linearity. The second block includes the acoustical channel and up link transducers which are assumed to be linear and have a comparatively longer impulse response and higher time variability. The second contribution is a new non-linear adaptive echo canceler which is based on the cascaded model and has greater robustness to changes in the acoustic channel than an existing power filter approach.
Moctar Mossi Idrissa, Christelle Yemdji, Nicholas W. D. Evans, Christophe Beaugeant, Philippe Degry
ICASSP3
2011 Handling overlaps in spoken term detection
abstract
Spoken term detection (STD) systems usually arrive at many overlapping detections which are often addressed with some pragmatic approaches, e.g. choosing the best detection to represent all the overlaps. In this paper we present a theoretical study based on a concept of acceptance space. In particular, we present two confidence estimation approaches based on Bayesian and evidence perspectives respectively. Analysis shows that both approaches possess respective ad vantages and shortcomings, and that their combination has the potential to provide an improved confidence estimation. Experiments conducted on meeting data confirm our analysis and show considerable performance improvement with the combined approach, in particular for out-of-vocabulary spoken term detection with stochastic pronunciation modeling.
Dong Wang 0013, Nicholas W. D. Evans, Raphaël Troncy, Simon King 0001
ICASSP2
2011 Semi-supervised face recognition with LDA self-training
abstract
Face recognition algorithms based on linear discriminant analysis (LDA) generally give satisfactory performance but tend to require a relatively high number of samples in order to learn reliable projections. In many practical applications of face recognition there is only a small number of labelled face images and in this case LDA-based algorithms generally lead to poor performance. The contributions in this paper relate to a new semi-supervised, self-training LDA-based algorithm which is used to augment a manually labelled training set with new data from an unlabelled, auxiliary set and hence to improve recognition performance. Without the cost of manual labelling such auxiliary data is often easily acquired but is not normally useful for learning. We report face recognition experiments on 3 independent databases which demonstrate a constant improvement of our baseline, supervised LDA system. The performance of our algorithm is also shown to significantly outperform other semi-supervised learning algorithms.
Xuran Zhao, Nicholas W. D. Evans, Jean-Luc Dugelay
ICIP2
2011 Online Pattern Learning for Non-Negative Convolutive Sparse Coding
abstract
The unsupervised learning of spectro-temporal speech patterns is relevant in a broad range of tasks. Convolutive non-negative matrix factorization (CNMF) and its sparse version, convolutive non-negative sparse coding (CNSC), are powerful, related tools. A particular difficulty of CNMF/CNSC, however, is the high demand on computing power and memory, which can prohibit their application to large scale tasks. In this paper, we propose an online algorithm for CNMF and CNSC, which processes input data piece-by-piece and updates the learned patterns after the processing of each piece by using accumulated sufficient statistics. The online CNSC algorithm remarkably increases converge speed of the CNMF/CNSC pattern learning, thereby enabling its application to large scale tasks.
Dong Wang 0013, Ravichander Vipperla, Nicholas W. D. Evans
INTERSPEECH3
2011 Parallel and Hierarchical Decision Making for Sparse Coding in Speech Recognition
abstract
Sparse coding exhibits promising performance in speech processing, mainly due to the large number of bases that can be used to represent speech signals. However, the high demand for computational power represents a major obstacle in the case of large datasets, as does the difficulty in utilising information scattered sparsely in high dimensional features. This paper reports the use of an online dictionary learning technique, proposed recently by the machine learning community, to learn large scale bases efficiently, and proposes a new parallel and hierarchical architecture to make use of the sparse information in high dimensional features. The approach uses multilayer perceptrons (MLPs) to model sparse feature subspaces and make local decisions accordingly; the latter are integrated by additional MLPs in a hierarchical way for making global decisions. Experiments on the WSJ database show that the proposed approach not only solves the problem of prohibitive computation with large-dimensional sparse features, but also provides better performance in a frame-level phone prediction task.
Dong Wang 0013, Ravichander Vipperla, Nicholas W. D. Evans
INTERSPEECH3
2010 The lia-eurecom RT'09 speaker diarization system: Enhancements in speaker modelling and cluster purification
abstract
There are two approaches to speaker diarization. They are bottom-up and top-down. Our work on top-down systems show that they can deliver competitive results compared to bottom-up systems and that they are extremely computationally efficient, but also that they are particularly prone to poor model initialisation and cluster impurities. In this paper we present enhancements to our state-of-the-art, top-down approach to speaker diarization that deliver improved stability across three different datasets composed of conference meetings from five standard NIST RT evaluations. We report an improved approach to speaker modelling which, despite having greater chances for cluster impurities, delivers a 35% relative improvement in DER for the MDM condition. We also describe new work to incorporate cluster purification into a top-down system which delivers relative improvements of 44% over the baseline system without compromising computational efficiency.
Simon Bozonnet, Nicholas W. D. Evans, Corinne Fredouille
ICASSP2
2010 An assessment of linear adaptive filter performance with nonlinear distortions
abstract
Acoustic echo cancellers are generally based on the assumption of a linear echo path between the transducers. However the small loudspeakers that are commonly used in todays terminals can introduce nonlinear distortions that reduce the performance of echo cancellation. In order to evaluate the degradation in performance, this paper assesses the behaviour of five linear echo cancellers in the presence of nonlinearities and presents the first thorough comparison of their robustness. Even if the performance of all the echo cancellers degrades as expected, some algorithms are shown to be more robust than others: fast converging algorithms and block signal processing are more perturbed in nonlinear environments.
Moctar Mossi Idrissa, Nicholas W. D. Evans, Christophe Beaugeant
ICASSP2
2010 System output combination for improved speaker diarization
abstract
International audience
Simon Bozonnet, Nicholas W. D. Evans, Xavier Anguera Miró, Oriol Vinyals, Gerald Friedland, Corinne Fredouille
INTERSPEECH2
2010 An integrated top-down/bottom-up approach to speaker diarization
abstract
International audience
Simon Bozonnet, Nicholas W. D. Evans, Corinne Fredouille, Dong Wang 0013, Raphaël Troncy
INTERSPEECH2
2010 CRF-based stochastic pronunciation modeling for out-of-vocabulary spoken term detection
abstract
Out-of-vocabulary (OOV) terms present a significant challenge to spoken term detection (STD). This challenge, to a large ex-tent, lies in the high degree of uncertainty in pronunciations of OOV terms. In previous work, we presented a stochastic pro-nunciation modeling (SPM) approach to compensate for this uncertainty. A shortcoming of our original work, however, is that the SPM was based on a joint-multigram model (JMM), which is suboptimal. In this paper, we propose to use con-ditional random fields (CRFs) for letter-to-sound conversion, which significantly improves quality of the predicted pronun-ciations. When applied to OOV STD, we achieve consider-able performance improvement with both a 1-best system and an SPM-based system. Index Terms: speech recognition, spoken term detection, con-ditional random field, joint multigram model
Dong Wang 0013, Simon King 0001, Nicholas W. D. Evans, Raphaël Troncy
INTERSPEECH3
2009 Speaker diarization using unsupervised discriminant analysis of inter-channel delay features
abstract
When multiple microphones are available estimates of inter-channel delay, which characterise a speaker's location, can be used as features for speaker diarization. Background noise and reverberation can, however, lead to noisy features and poor performance. To ameliorate these problems, this paper presents a new approach to the discriminant analysis of delay features for speaker diarization. This novel and nonetheless unsupervised approach aims to increase speaker separability in delay-space. We assess the approach on subsets of four standard NIST RT datasets and demonstrate a relative improvement in diarization error rate of 25% on a separate evaluation set using delay features alone.
Nicholas W. D. Evans, Corinne Fredouille, Jean-François Bonastre
ICASSP1
2009 Towards a New Image-Based Spectrogram Segmentation Speech Coder Optimised for Intelligibility
Keith A. Jellyman, Nicholas W. D. Evans, Wei Ming Liu 0002, John S. D. Mason
MMM2
2008 A large scale footstep database for biometric studies created using cross-biometrics for labelling
abstract
This paper describes a semi-automatic system to capture and label a reasonable size biometric database. In our case, the biometric to be assessed are footstep signals, but the system could be extendable to other biometrics. Extra biometric data such as the voice and video recordings of the face and the gait are used to assist the database labelling to minimise the error. Thus, audio identifier recordings are used to automatically label the database with a speaker recognition system achieving results of 0.15% of equal error rate (EER) of person verification using Gaussian mixture models (GMM). Also, a footstep detector system has been developed to reduce the presence of invalid signals from the database having a percentage of less than 1% of correct footsteps miss-classified using features from the ground reaction force (GRF) and using a support vector machine (SVM) classifier. To date, more than 20,000 footstep signals have been collected from more than 100 people, which is well beyond previously reported databases. The database is collected in different sessions which will allow us to study how different factors such as footwear, the person carrying a load or different walking speeds affect the recognition of persons using their footsteps.
Rubén Vera-Rodríguez, Richard P. Lewis 0002, John S. D. Mason, Nicholas W. D. Evans
ICARCV4
2008 Assessment of objective quality measures for speech intelligibility
abstract
This paper assesses 9 prominent objective quality measures for their potential in intelligibility estimation. Degradation considered include additive noises and those introduced by coding and enhancement schemes, totalling 78 types. This paper is believed to be the first to conduct an assessment on such a large combination of quality measures and degradations allowing side-byside analysis. Experimental results show that the sophisticated perceptual-based measures which are superior for quality estimation, do not necessarily correlate well with human intelligibility and, in fact, give poorer correlations when enhancement schemes are considered. Meanwhile, the weighted spectral slope (WSS) emerges to be the most promising approach among all measures considered, scoring the highest correlation in 5 out of the 6 test sets. Worth noting are the positive correlations obtained with WSS which range from 0.14 to 0.86, as opposed to those with PESQ from -0.58 to 0.74. Such findings put WSS, a relatively conventional measure, in a new light as a potential intelligibility assessor.
Wei Ming Liu 0002, Keith A. Jellyman, Nicholas W. D. Evans, John S. D. Mason
INTERSPEECH3
2007 Influence of task duration in text-independent speaker verification
abstract
International audience
Benoit G. B. Fauve, Nicholas W. D. Evans, Neil Pearson, Jean-François Bonastre, John S. D. Mason
INTERSPEECH2
2007 The influence of speech activity detection and overlap on speaker diarization for meeting room recordings
abstract
Abstract This paper addresses the problem of speaker diarization inthe specific context of meeting room recordings which of-ten involve a high degree of spontaneous speech with largeoverlapped speech segments, speaker noise (laughs, whispers,coughs, etc.) and very short speaker turns. A large variabilityin signal quality has brought an additional level of complexity.This paper investigates the effects of speech activity detectionand overlapped speech through speaker diarization experimentsconducted on the NIST RT’05 and RT’06 data sets. Resultsindicate that our system is highly sensitive to the shape of theinitial segmentation and that, perhaps surprisingly, perfect ref-erences can even degrade performance. Finally we propose adirection for future research to incorporate confidence valuesaccording to acoustic attributes in order to unify what is cur-rently a somewhat disjointed approach to speaker diarization. Index Terms : speaker diarization, meeting room, speech activ-ity detection, overlapped speech.
Corinne Fredouille, Nicholas W. D. Evans
INTERSPEECH2
2006 An Assessment on the Fundamental Limitations of Spectral Subtraction
abstract
As with many approaches to noise robust automatic speech recognition (ASR) the benefits of spectral subtraction tend to diminish as noise levels in the order of 0 dB are approached. Whilst the majority of related work focuses on reducing magnitude errors a number of new approaches addressing the often overlooked, additional sources of error have appeared in the literature in recent years. Relatively lacking in the literature, however, is an empirical assessment which compares the effects of each error when noisy speech is processed by spectral subtraction. Such studies are vital in order to appreciate the potential penalty in performance when sources of error are overlooked. The objective in this paper is to assess, through ASR, the performance penalty associated with each source of error when noisy speech is treated with spectral subtraction. Experimental evidence based on two standard European databases and ASR protocols illustrates that, perhaps contrary to popular belief, for noise levels in the order of 0 dB and below, these often overlooked sources of error can lead to non-negligible degradations in performance. Whilst not a new idea, here the original emphasis is a thorough assessment that empirically highlights both the fundamental limitations and potential benefit of including the full complement of errors in the spectral subtraction model
Nicholas W. D. Evans, John S. D. Mason, Wei Ming Liu 0002, Benoit G. B. Fauve
ICASSP (1)1
2006 Assessment of Objective Quality Measures for Speech Intelligibility Estimation
abstract
This paper investigates the accuracy of automatic speech recognition (ASR) and 6 other well-reported objective quality measures for the task of estimating speech intelligibility. It is believed to be the first assessment of such a range of measures side-by-side and in the context of intelligibility. A total of 39 degradation conditions including those from a newly proposed low bit rate (0.3 to 1.5 kbps) codec and a noise suppression system are considered. They provide real and varied scenarios to assess the measures. The objective scores are compared to subjective listening scores, and their correlation used to assess the approach. All tests are conducted on the European standard Aurora 2 corpus. Experiments show that ASR and perceptual estimation of speech quality (PESQ) are potentially reliable estimators of intelligibility with subjective correlation as high as 0.99 and 0.96 respectively. Furthermore, ASR gives a trend corresponding to that of subjective intelligibility assessment for the different configurations of the new codec, while most others fail
Wei Ming Liu 0002, Keith A. Jellyman, John S. D. Mason, Nicholas W. D. Evans
ICASSP (1)4
2006 An assessment of automatic speech recognition as speech intelligibility estimation in the context of additive noise
abstract
This paper investigates the potential applicability of automatic speech recognition (ASR) and 6 well-reported objective quality measures for the task of ranking intelligibility of speech degraded by different real life background noises. In a recent investigation ASR has been reported to give high subjective correlation with human assessment when tested with various system degradations. This paper extends this investigation in two directions. First, the usefulness of the measures in the context of different real-life noises is considered. Second, the direct correspondence between statistics computed by an ASR system and human perceived intelligibility is assessed. Subjective listening tests are carried out to provide ground truth. Results show that ASR and WSS (weighted spectral slope) are the only two measures out of the seven considered to give good correlation with human opinion. Specially noted is performance of ASR with correlations ranging from 0.77 to 0.90. Index Terms: speech intelligibility, objective quality measures, ASR
Wei Ming Liu 0002, John S. D. Mason, Nicholas W. D. Evans, Keith A. Jellyman
INTERSPEECH3
2003 Morphological filtering of speech spectrograms in the context of additive noise
Francisco Romero Rodriguez, Wei Ming Liu 0002, Nicholas W. D. Evans, John S. D. Mason
INTERSPEECH3
2002 Computationally efficient noise compensation for robust automatic speech recognition assessed under the Aurora 2/3 framework
abstract
In the context of mobile telephony there is a need for low resource, computationally efficient noise compensation and speech enhancement approaches. This paper assesses the performance of efficient quantile-based noise estimation integrated into a nonlinear spectral subtraction framework. The approach has been implemented in real-time with minimal latency on a 500Mhz processor and is well within the processing capabilities. Experiments are reported on the AURORA 2 and AURORA 3 corpa. Results show an average relative improvement of 15% on the clean and multicondition training sets of the AURORA 2 database and an overall average relative improvement of 20% across the four AURORA 3 databases. It is acknowledged that these are not state-of-the-art results and further optimisation is anticipated.
Nicholas W. D. Evans, John S. D. Mason
INTERSPEECH1
2001 Noise estimation without explicit speech, non-speech detection: a comparison of mean, modal and median based approaches
abstract
Automatic speech recognition performance tends to be degraded in noisy conditions. Spectral subtraction is a simple, popular approach of noise compensation. In conventional spectral subtraction [1, 2], noise statistics are updated during speech gaps and subtracted from a corrupt signal during speech intervals. Some means of explicit speech, non-speech detection is therefore essential. Recent proposals have avoided the problem of speech, non-speech detection [3, 4, 5, 6, 7] by continually updating noise estimates whether speech is present or not.
Nicholas W. D. Evans, John S. D. Mason
INTERSPEECH1