EDBT 2026 Demo / reviewers in the wild / expert
Anthony Larcher
dblp:44/5405
· DBLP profile ↗
45ranked-venue papers
13as first author
15since 2021 · last 2026
0000-0003-4398-0224ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 13 first-author · 10 since 2021Artificial intelligence and machine learning · 34 · 5 first-author · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SENS-ASR: Semantic Embedding Injection in Neural-transducer for Streaming Automatic Speech Recognition
Youness Dkhissi, Valentin Vielzeuf, Elys Allesiardo, Anthony Larcher |
LREC | 4 |
| 2025 | When The MOS Predictor Asks For Training Annotation In Cross Lingual/Domain Adaptation
Natacha Miniconi, Meysam Shamsi, Anthony Larcher |
INTERSPEECH | 3 |
| 2024 | ALLIES: A Speech Corpus for Segmentation, Speaker Diarization, Speech Recognition and Speaker Change DetectionabstractThis paper presents ALLIES, a meta corpus which gathers and extends existing French corpora collected from radio and TV shows. The corpus contains 1048 audio files for about 500 hours of speech. Agglomeration of data is always a difficult issue, as the guidelines used to collect, annotate and transcribe speech are generally different from one corpus to another. ALLIES intends to homogenize and correct speaker labels among the different files by integrated human feedback within a speaker verification system. The main contribution of this article is the design of a protocol in order to evaluate properly speech segmentation (including music and overlap detection), speaker diarization, speech transcription and speaker change detection. As part of it, a test partition has been carefully manually 1) segmented and annotated according to speech, music, noise, speaker labels with specific guidelines for overlap speech, 2) orthographically transcribed. This article also provides as a second contribution baseline results for several speech processing tasks. Marie Tahon, Anthony Larcher, Martin Lebourdais, Fethi Bougares, Anna Silnova, Pablo Gimeno |
LREC/COLING | 2 |
| 2024 | ASoBO: Attentive Beamformer Selection for Distant Speaker Diarization in MeetingsabstractSpeaker Diarization (SD) aims at grouping speech segments that belong to the same speaker. This task is required in many speech-processing applications, such as rich meeting transcription. In this context, distant microphone arrays usually capture the audio signal. Beamforming, i.e., spatial filtering, is a common practice to process multi-microphone audio data. However, it often requires an explicit localization of the active source to steer the filter. This paper proposes a self-attention-based algorithm to select the output of a bank of fixed spatial filters. This method serves as a feature extractor for joint Voice Activity (VAD) and Overlapped Speech Detection (OSD). The speaker diarization is then inferred from the detected segments. The approach shows convincing distant VAD, OSD, and SD performance, e.g. 14.5% DER on the AISHELL-4 dataset. The analysis of the self-attention weights demonstrates their explainability, as they correlate with the speaker's angular locations. Théo Mariotte, Anthony Larcher, Silvio Montrésor, Jean-Hugh Thomas |
INTERSPEECH | 2 |
| 2024 | Channel-Combination Algorithms for Robust Distant Voice Activity and Overlapped Speech DetectionabstractVoice Activity Detection (VAD) and Overlapped Speech Detection (OSD) are key pre-processing tasks for speaker diarization. In the meeting context, it is often easier to capture speech with a distant device. This consideration however leads to severe performance degradation. We study a unified supervised learning framework to solve distant multi-microphone joint VAD and OSD (VAD+OSD). This paper investigates various multi-channel VAD+OSD front-ends that weight and combine incoming channels. We propose three algorithms based on the Self-Attention Channel Combinator (SACC), previously proposed in the literature. Experiments conducted on the AMI meeting corpus exhibit that channel combination approaches bring significant VAD+OSD improvements in the distant speech scenario. Specifically, we explore the use of learned complex combination weights and demonstrate the benefits of such an approach in terms of explainability. Channel combination-based VAD+OSD systems are evaluated on the final back-end task, i.e. speaker diarization, and show significant improvements. Finally, since multi-channel systems are trained given a fixed array configuration, they may fail in generalizing to other array set-ups, e.g. mismatched number of microphones. A channel-number invariant loss is proposed to learn a unique feature representation regardless of the number of available microphones. The evaluation conducted on mismatched array configurations highlights the robustness of this training strategy. Théo Mariotte, Anthony Larcher, Silvio Montrésor, Jean-Hugh Thomas |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Multi-microphone Automatic Speech Segmentation in Meetings Based on Circular Harmonics FeaturesabstractSpeaker diarization is the task of answering Who spoke and when? in an audio stream.Pipeline systems rely on speech segmentation to extract speakers' segments and achieve robust speaker diarization.This paper proposes a common framework to solve three segmentation tasks in the distant speech scenario: Voice Activity Detection (VAD), Overlapped Speech Detection (OSD), and Speaker Change Detection (SCD).In the literature, a few studies investigate the multi-microphone distant speech scenario.In this work, we propose a new set of spatial features based on direction-of-arrival estimations in the circular harmonic domain (CH-DOA).These spatial features are extracted from multi-microphone audio data and combined with standard acoustic features.Experiments on the AMI meeting corpus show that CH-DOA can improve the segmentation while being robust in case of deactivated microphones. Théo Mariotte, Anthony Larcher, Silvio Montrésor, Jean-Hugh Thomas |
INTERSPEECH | 2 |
| 2023 | Towards lifelong human assisted speaker diarization
Meysam Shamsi, Anthony Larcher, Loïc Barrault, Sylvain Meignier, Yevhenii Prokopalo, Marie Tahon, Ambuj Mehrish, Simon Petit-Renaud, Olivier Galibert, Samuel Gaist, André Anjos, Sébastien Marcel, Marta R. Costa-jussà |
Comput. Speech Lang. | 2 |
| 2022 | Are disentangled representations all you need to build speaker anonymization systems?abstractInternational audience Pierre Champion, Anthony Larcher, Denis Jouvet |
INTERSPEECH | 2 |
| 2022 | Microphone Array Channel Combination Algorithms for Overlapped Speech DetectionabstractInternational audience Théo Mariotte, Anthony Larcher, Silvio Montrésor, Jean-Hugh Thomas |
INTERSPEECH | 2 |
| 2022 | Overlaps and Gender Analysis in the Context of Broadcast MediaabstractOur main goal is to study the interactions between speakers according to their gender and role in broadcast media. In this paper, we propose an extensive study of gender and overlap annotations in various speech corpora mainly dedicated to diarisation or transcription tasks. We point out the issue of the heterogeneity of the annotation guidelines for both overlapping speech and gender categories. On top of that, we analyse how the speech content (casual speech, meetings, debate, interviews, etc.) impacts the distribution of overlapping speech segments. On a small dataset of 93 recordings from LCP French channel, we intend to characterise the interactions between speakers according to their gender. Finally, we propose a method which aims to highlight active speech areas in terms of interactions between speakers. Such a visualisation tool could improve the efficiency of qualitative studies conducted by researchers in human sciences. Martin Lebourdais, Marie Tahon, Antoine Laurent, Sylvain Meignier, Anthony Larcher |
LREC | 5 |
| 2021 | On the Invertibility of a Voice Privacy System Using Embedding AlignmentabstractThis paper explores various attack scenarios on a voice anonymization system using embeddings alignment techniques. We use Wasserstein-Procrustes (an algorithm initially designed for unsupervised translation) or Procrustes analysis to match two sets of$x$-vectors, before and after voice anonymization, to mimic this transformation as a rotation function. We compute the optimal rotation and compare the results of this approximation to the official Voice Privacy Challenge results. We show that a complex system like the baseline of the Voice Privacy Challenge can be approximated by a rotation, estimated using a limited set of$x$-vectors. This paper studies the space of solutions for voice anonymization within the specific scope of rotations. Rotations being reversible, the proposed method can recover up to 62% of the speaker identities from anonymized embeddings. Pierre Champion, Thomas Thebaud, Gaël Le Lan, Anthony Larcher, Denis Jouvet |
ASRU | 4 |
| 2021 | Speaker Embeddings for Diarization of Broadcast Data In The Allies ChallengeabstractDiarization consists in the segmentation of speech signals and the clustering of homogeneous speaker segments. State-of-the-art systems typically operate upon speaker embeddings, such as i-vectors or neural x-vectors, extracted from mel cepstral coefficients (MFCCs) or spectrograms. The recent SincNet architecture extracts x-vectors directly from raw speech signals. The work reported in this paper compares the performance of different embeddings extracted from MFCCs or the raw signal for speaker diarization and broadcast media treated with compression and sub-sampling, operations which typically degrade performance. Experiments are performed with the new ALLIES database that was designed to complement existing, publicly available French corpora of broadcast radio and TV shows. Results show that, in adverse conditions, with compression and sampling mismatch, SincNet x-vectors outperform i-vectors and x-vectors by relative DERs of 43% and 73% respectively. Additionally we found that SincNet x-vectors are not the absolute best embeddings but are more robust to data mismatch than others. Anthony Larcher, Ambuj Mehrish, Marie Tahon, Sylvain Meignier, Jean Carrive, David Doukhan, Olivier Galibert, Nicholas W. D. Evans |
ICASSP | 1 |
| 2021 | End-to-End anti-spoofing with RawNet2abstractSpoofing countermeasures aim to protect automatic speaker verification systems from being manipulated by spoofed speech signals. While results from the most recent ASVspoof 2019 evaluation show great potential to detect most forms of attack, some continue to evade detection. This paper reports the first application of RawNet2 to anti-spoofing. RawNet2 ingests raw audio and has potential to learn cues that are not detectable using more traditional countermeasure solutions. We describe modifications made to the original RawNet2 architecture so that it can be applied to anti-spoofing. For A17 attacks, our RawNet2 systems results are the second-best reported, while the fusion of RawNet2 and baseline countermeasures gives the second-best results reported for the full ASVspoof 2019 logical access condition. Our results are reproducible with open source software. Hemlata Tak, Jose Patino 0001, Massimiliano Todisco, Andreas Nautsch, Nicholas W. D. Evans, Anthony Larcher |
ICASSP | 6 |
| 2021 | Handwritten Digits Reconstruction from Unlabelled EmbeddingsabstractIn this paper, we investigate template reconstruction attack of touchscreen biometrics, based on handwritten digits writer verification. In the event of a template database theft, we show that reconstructing the original drawn digit from the embeddings is possible without access to the original embedding encoder. Using an external labelled dataset, an attack encoder is trained along with a Mixture Density Recurrent Neural Network decoder. Thanks to an alignment flow, initialized with Linear Discriminant Analysis and Procrustes, the transfer function between the output space of the original and the attack encoder is estimated. The successive application of transfer function and decoder to the stolen embeddings allows to reconstruct the original drawings, which can be used to spoof the behavioural biometrics system. Thomas Thebaud, Gaël Le Lan, Anthony Larcher |
ICASSP | 3 |
| 2021 | The LIUM Human Active Correction Platform for Speaker Diarization
Alexandre Flucha, Anthony Larcher, Ambuj Mehrish, Sylvain Meignier, Florian Plaut, Nicolas Poupon, Yevhenii Prokopalo, Adrien Puertolas, Meysam Shamsi, Marie Tahon |
Interspeech | 2 |
| 2020 | Evaluation of Lifelong Learning SystemsabstractCurrent intelligent systems need the expensive support of machine learning experts to sustain their performance level when used on a daily basis. To reduce this cost, i.e. remaining free from any machine learning expert, it is reasonable to implement lifelong (or continuous) learning intelligent systems that will continuously adapt their model when facing changing execution conditions. In this work, the systems are allowed to refer to human domain experts who can provide the system with relevant knowledge about the task. Nowadays, the fast growth of lifelong learning systems development rises the question of their evaluation. In this article we propose a generic evaluation methodology for the specific case of lifelong learning systems. Two steps will be considered. First, the evaluation of human-assisted learning (including active and/or interactive learning) outside the context of lifelong learning. Second, the system evaluation across time, with propositions of how a lifelong learning intelligent system should be evaluated when including human assisted learning or not. Yevhenii Prokopalo, Sylvain Meignier, Olivier Galibert, Loïc Barrault, Anthony Larcher |
LREC | 5 |
| 2020 | Introduction to the special issue "Speaker and language characterization and recognition: Voice modeling, conversion, synthesis and ethical aspects"
Jean-François Bonastre, Tomi Kinnunen, Anthony Larcher, Junichi Yamagishi |
Comput. Speech Lang. | 3 |
| 2019 | I4U Submission to NIST SRE 2018: Leveraging from a Decade of Shared ExperiencesabstractThe I4U consortium was established to facilitate a joint entry to NIST speaker recognition evaluations (SRE). The latest edition of such joint submission was in SRE 2018, in which the I4U submission was among the best-performing systems. SRE'18 also marks the 10-year anniversary of I4U consortium into NIST SRE series of evaluation. The primary objective of the current paper is to summarize the results and lessons learned based on the twelve sub-systems and their fusion submitted to SRE'18. It is also our intention to present a shared view on the advancements, progresses, and major paradigm shifts that we have witnessed as an SRE participant in the past decade from SRE'08 to SRE'18. In this regard, we have seen, among others, a paradigm shift from supervector representation to deep speaker embedding, and a switch of research challenge from channel compensation to domain adaptation. Kong-Aik Lee, Ville Hautamäki, Tomi Kinnunen, Hitoshi Yamamoto, Koji Okabe, Ville Vestman, Jing Huang 0019, Guo-Hong Ding, Hanwu Sun, Anthony Larcher, Rohan Kumar Das, Haizhou Li 0001, Mickael Rouvier, Pierre-Michel Bousquet, Wei Rao 0002, Qing Wang 0039, Fahimeh Bahmaninezhad, Héctor Delgado, Massimiliano Todisco |
INTERSPEECH | 10 |
| 2018 | An Open-Source Speaker Gender Detection Framework for Monitoring Gender EqualityabstractThis paper presents an approach based on acoustic analysis to describe gender equality in French audiovisual streams, through the estimation of male and female speaking time. Gender detection systems based on Gaussian Mixture Models, i-vectors and Convolutional Neural Networks (CNN) were trained using an internal database of 2,284 French speakers and evaluated using REPERE challenge corpus. The CNN system obtained the best performance with a frame-level gender detection F-measure of 96.52 and a hourly women speaking time percentage error bellow 0.6%. It was considered reliable enough to realize large-scale gender equality descriptions. The proposed gender detection system has been packaged as an open-source framework. David Doukhan, Jean Carrive, Félicien Vallet, Anthony Larcher, Sylvain Meignier |
ICASSP | 4 |
| 2018 | S4D: Speaker Diarization Toolkit in PythonabstractInternational audience Pierre-Alexandre Broux, Florent Desnous, Anthony Larcher, Simon Petit-Renaud, Jean Carrive, Sylvain Meignier |
INTERSPEECH | 3 |
| 2018 | An Adaptive Method for Cross-Recording Speaker DiarizationabstractNowadays, state-of-the-art speaker diarization systems heavily rely on between-recording variability compensation methods to accurately process large collections of recordings. Variability estimation is performed on consequent training datasets, which must be labeled by speaker. One major problem of such systems is the acoustic mismatch between training and target data that degrades performances. Most of the collections contain lots of speakers speaking in various acoustic conditions. In this paper, we investigate how unlabeled speakers can help improve between-recording variability estimation, to overcome the mismatch issue. We propose a scalable unsupervised adaptation framework for two types of variability compensation. The proposed framework consists in adapting a state-of-the-art diarization and linking system, trained on out-of-domain data, using the data of the collection itself. Experiments in mismatch condition are run on two French Television shows, while the initial training dataset is composed of Radio recordings. Results indicate that the proposed adaptation framework reduces the cross-recording DER of 13% in average for variable collection sizes. Gaël Le Lan, Delphine Charlet, Anthony Larcher, Sylvain Meignier |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | A Triplet Ranking-Based Neural Network for Speaker Diarization and LinkingabstractInternational audience Gaël Le Lan, Delphine Charlet, Anthony Larcher, Sylvain Meignier |
INTERSPEECH | 3 |
| 2017 | The I4U Mega Fusion and Collaboration for NIST Speaker Recognition Evaluation 2016abstract18th Annual Conference of the International Speech Communication Association, INTERSPEECH 2017, Stockholm, Sweden, 20-24 August 2017 Kong-Aik Lee, Ville Hautamäki, Tomi Kinnunen, Anthony Larcher, Andreas Nautsch, Themos Stafylakis, Gang Liu 0001, Mickael Rouvier, Wei Rao 0002, Federico Alegre, Man-Wai Mak, Achintya Kumar Sarkar, Héctor Delgado, Rahim Saeidi, Hagai Aronowitz, Aleksandr Sizov, Hanwu Sun, Trung Hieu Nguyen 0001, Guangsen Wang, Bin Ma 0001, Ville Vestman, Md. Sahidullah, M. Halonen, Anssi Kanervisto, Gaël Le Lan, Fahimeh Bahmaninezhad, Sergey Isadskiy, Christian Rathgeb, Christoph Busch 0001, Georgios Tzimiropoulos, Q. Qian, Q. Zhao, J. Xue, R. Jin, T. Zhao, Pierre-Michel Bousquet, Moez Ajili, Waad Ben Kheder, Driss Matrouf, Zhi Hao Lim, Chenglin Xu, Haihua Xu 0001, Chng Eng Siong, Benoit G. B. Fauve, Kaavya Sriskandaraja, Vidhyasaharan Sethu, W. W. Lin, Dennis Alexander Lehmann Thomsen, Zheng-Hua Tan, Massimiliano Todisco, Nicholas W. D. Evans, Haizhou Li 0001, John H. L. Hansen, Jean-François Bonastre, Eliathamby Ambikairajah |
INTERSPEECH | 4 |
| 2016 | An extensible speaker identification sidekit in PythonabstractSIDEKIT is a new open-source Python toolkit that includes a large panel of state-of-the-art components and allow a rapid prototyping of an end-to-end speaker recognition system. For each step from front-end feature extraction, normalization, speech activity detection, modelling, scoring and visualization, SIDEKIT offers a wide range of standard algorithms and flexible interfaces. The use of a single efficient programming and scripting language (Python in this case), and the limited dependencies, facilitate the deployment for industrial applications and extension to include new algorithms as part of the whole tool-chain provided by SIDEKIT. Performance of SIDEKIT is demonstrated on two standard evaluation tasks, namely the RSR2015 and NIST-SRE 2010. Anthony Larcher, Kong-Aik Lee, Sylvain Meignier |
ICASSP | 1 |
| 2016 | Iterative PLDA Adaptation for Speaker DiarizationabstractInternational audience Gaël Le Lan, Delphine Charlet, Anthony Larcher, Sylvain Meignier |
INTERSPEECH | 3 |
| 2016 | The 2015 NIST Language Recognition Evaluation: The Shared View of I2R, Fantastic4 and SingaMSabstractTechnical report for NIST LRE 2015 Workshop Kong-Aik Lee, Haizhou Li 0001, Li Deng 0001, Ville Hautamäki, Wei Rao 0002, Anthony Larcher, Hanwu Sun, Trung Hieu Nguyen 0001, Guangsen Wang, Aleksandr Sizov, Jianshu Chen, Ivan Kukanov, Amir Hossein Poorjam, Trung Ngo Trong, Chenglin Xu, Haihua Xu 0001, Bin Ma 0001, Chng Eng Siong, Sylvain Meignier |
INTERSPEECH | 7 |
| 2015 | The reddots data collection for speaker recognitionabstractde niveau recherche, publiés ou non, émanant des établissements d'enseignement et de recherche français ou étrangers, des laboratoires publics ou privés. Kong-Aik Lee, Anthony Larcher, Guangsen Wang, Patrick Kenny, Niko Brümmer, David A. van Leeuwen, Hagai Aronowitz, Marcel Kockmann, Carlos Vaquero, Bin Ma 0001, Haizhou Li 0001, Themos Stafylakis, Jahangir Alam 0001, Albert Swart, Javier Perez |
INTERSPEECH | 2 |
| 2014 | Modelling the alternative hypothesis for text-dependent speaker verificationabstractThis paper describes text-dependent speaker verification as a task involving four classes of trials depending on whether the target speaker or an impostor pronounces the expected pass-phrase or not. These four classes are used to reformulate the log-likelihood ratio traditionally used in text-independent speaker verification. Three formulations of the alternative hypothesis are considered, leading to three new expressions of the verification score. Experiments performed on the publicly available RSR2015 database show a significant improvement compared to existing baseline scores. A relative gain up to 61% in term of minimum cost is achieved when considering that the alternative hypothesis is the union of three sub-hypotheses corresponding to the three existing classes of impostures. Anthony Larcher, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001 |
ICASSP | 1 |
| 2014 | Imposture classification for text-dependent speaker verificationabstractThis work focuses on text-dependent speaker verification, where a user is required to chose and pronounce a customized pass-phrase to get authenticated. In this context, there are three types of impostures: an impostor pronouncing the correct pass-phrase, an impostor pronouncing a wrong pass-phrase and the most difficult one: an impostor playing back a recording of the target speaker pronouncing a wrong pass-phrase. Detecting and classifying different types of impostures can help to prevent future impostures of the same type. In this work, we first propose a new verification score to reject Playback impostures. This score allows a relative reduction of 90% of the equal error rate against Playback impostures while offering performance similar to the baseline text-dependent score against other types of impostures. As a second contribution, we show that the new score can be combined with an existing text-dependent verification score to improve the classification of the different types of impostures. The performance of the speaker verification engine for imposture classification is significantly improved with the Cllrdecreasing by at least 29% compared to the original system. Anthony Larcher, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001 |
ICASSP | 1 |
| 2014 | Extended RSR2015 for text-dependent speaker verification over VHF channelabstractInternational audience Anthony Larcher, Kong-Aik Lee, Pablo Luis Sordo Martinez, Trung Hieu Nguyen 0001, Bin Ma 0001, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2014 | Text-dependent speaker verification: Classifiers, databases and RSR2015abstractThe RSR2015 database, designed to evaluate text-dependent speaker verification systems under different durations and lexical constraints has been collected and released by the Human Language Technology (HLT) department at Institute for Infocomm Research (I2R) in Singapore. English speakers were recorded with a balanced diversity of accents commonly found in Singapore. More than 151 h of speech data were recorded using mobile devices. The pool of speakers consists of 300 participants (143 female and 157 male speakers) between 17 and 42 years old making the RSR2015 database one of the largest publicly available database targeted for text-dependent speaker verification. We provide evaluation protocol for each of the three parts of the database, together with the results of two speaker verification system: the HiLAM system, based on a three layer acoustic architecture, and an i-vector/PLDA system. We thus provide a reference evaluation scheme and a reference performance on RSR2015 database to the research community. The HiLAM outperforms the state-of-the-art i-vector system in most of the scenarios. Anthony Larcher, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001 |
Speech Commun. | 1 |
| 2013 | Phonetically-constrained PLDA modeling for text-dependent speaker verification with multiple short utterancesabstractThe importance of phonetic variability for short duration speaker verification is widely acknowledged. This paper assesses the performance of Probabilistic Linear Discriminant Analysis (PLDA) and i-vector normalization for a text-dependent verification task. We show that using a class definition based on both speaker and phonetic content significantly improves the performance of a state-of-the-art system. We also compare four models for computing the verification scores using multiple enrollment utterances and show that using PLDA intrinsic scoring obtains the best performance in this context. This study suggests that such scoring regime remains to be optimized. Anthony Larcher, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001 |
ICASSP | 1 |
| 2013 | Automatic regularization of cross-entropy cost for speaker recognition fusionabstract\n Contains fulltext :\n 116325.pdf (author's version ) (Open Access)\n Ville Hautamäki, Kong-Aik Lee, David A. van Leeuwen, Rahim Saeidi, Anthony Larcher, Tomi Kinnunen, Taufiq Hasan, Seyed Omid Sadjadi, Gang Liu 0001, Hynek Boril, John H. L. Hansen, Benoit G. B. Fauve |
INTERSPEECH | 5 |
| 2013 | ALIZE 3.0 - open source toolkit for state-of-the-art speaker recognitionabstractInternational audience Anthony Larcher, Jean-François Bonastre, Benoit G. B. Fauve, Kong-Aik Lee, Christophe Lévy, Haizhou Li 0001, John S. D. Mason, Jean-Yves Parfait |
INTERSPEECH | 1 |
| 2013 | Multi-session PLDA scoring of i-vector for partially open-set speaker detectionabstractInternational audience Kong-Aik Lee, Anthony Larcher, Chang Huai You, Bin Ma 0001, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2013 | Vulnerability evaluation of speaker verification under voice conversion spoofing: the effect of text constraintsabstractVoice conversion, a technique to change one's voice to sound like that of another, poses a threat to even high performance speaker verification system. Vulnerability of text-independent speaker verification systems under spoofing attack, using statistical voice conversion technique, was evaluated and confirmed in our previous work. In this paper, we further extend the study to text-dependent speaker verification systems. In particular, we compare both joint density Gaussian mixture model (JD-GMM) and unit-selection (US) spoofing methods and, for the first time, the performances of text-independent and text-dependent speaker verification systems in a single study. We conduct the experiments using RSR2015 database which is recorded using multiple mobile devices. The experimental results indicate that text-dependent speaker verification system tolerates spoofing attacks better than the text-independent counterpart. Zhizheng Wu 0001, Anthony Larcher, Kong-Aik Lee, Chng Eng Siong, Tomi Kinnunen, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2013 | I4u submission to NIST SRE 2012: a large-scale collaborative effort for noise-robust speaker verificationabstractI4U is a joint entry of nine research Institutes and Universities across 4 continents to NIST SRE 2012. It started with a brief discussion during the Odyssey 2012 workshop in Singapore. An online discussion group was soon set up, providing a discussion platform for different issues surrounding NIST SRE’12. Noisy test segments, uneven multi-session training, variable enrollment duration, and the issue of open-set identification were actively discussed leading to various solutions integrated to the I4U submission. The joint submission and several of its 17 sub-systems were among top-performing systems. We summarize the lessons learnt from this large-scale effort. Rahim Saeidi, Kong-Aik Lee, Tomi Kinnunen, Tawfik Hasan, Benoit G. B. Fauve, Pierre-Michel Bousquet, Elie Khoury 0001, Pablo Luis Sordo Martinez, Jia Min Karen Kua, Chang Huai You, Hanwu Sun, Anthony Larcher, Padmanabhan Rajan, Ville Hautamäki, Cemal Hanilçi, Billy Braithwaite, Rosa González Hautamäki, Seyed Omid Sadjadi, Gang Liu 0001, Hynek Boril, Navid Shokouhi, Driss Matrouf, Laurent El Shafey, Pejman Mowlaee, Julien Epps, Tharmarajah Thiruvaran, David A. van Leeuwen, Bin Ma 0001, Haizhou Li 0001, John H. L. Hansen, Jean-François Bonastre, Sébastien Marcel, John S. D. Mason, Eliathamby Ambikairajah |
INTERSPEECH | 12 |
| 2012 | I-vectors in the context of phonetically-constrained short utterances for speaker verificationabstractShort speech duration remains a critical factor of performance degradation when deploying a speaker verification system. To overcome this difficulty, a large number of commercial applications impose the use of fixed pass-phrases. In this context, we show that the performance of the popular i-vector approach can be greatly improved by taking advantage of the phonetic information that they convey. Moreover, as i-vectors require a conditioning process to reach high accuracy, we show that further improvements are possible by taking advantage of this phonetic information within the normalisation process. We compare two methods, Within Class Covariance Normalization (WCCN) and Eigen Factor Radial (EFR), both relying on parameters estimated on the same development data. Our study suggests that WCCN is more robust to data mismatch but less efficient than EFR when the development data has a better match with the test data. Anthony Larcher, Pierre-Michel Bousquet, Kong-Aik Lee, Driss Matrouf, Haizhou Li 0001, Jean-François Bonastre |
ICASSP | 1 |
| 2012 | PLDA Modeling in I-Vector and Supervector Space for Speaker VerificationabstractIn this paper, we advocate the use of uncompressed form of i-vector. We employ the probabilistic linear discriminant analysis (PLDA) to handle speaker and session variability for speaker verification task. An i-vector is a low-dimensional vector containing both speaker and channel information acquired from a speech segment. When PLDA is used on i-vector, dimension reduction is performed twice – first in the i-vector extraction process and second in the PLDA model. Keeping the full dimensionality of i-vector in the supervector space for PLDA modeling and scoring would avoid unnecessary loss of information. The drawback of using PLDA on uncompressed i-vector is the inversion of large matrices, which we show can be solved rather efficiently by portioning large matrix into smaller blocks. We also introduce the Gaussianized rank-norm, as an alternative to whitening, for feature normalization prior to PLDA modeling. Index Terms: speaker verification, i-vector, probabilistic LDA 1. Kong-Aik Lee, Zhenmin Tang, Bin Ma 0001, Anthony Larcher, Haizhou Li 0001 |
INTERSPEECH | 5 |
| 2012 | RSR2015: Database for Text-Dependent Speaker Verification using Multiple Pass-PhrasesabstractInternational audience Anthony Larcher, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2011 | Joint Application of Speech and Speaker Recognition for Automation and Security in Smart Home
Kong-Aik Lee, Anthony Larcher, Helen Thai, Bin Ma 0001, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2011 | Spoken Language Recognition in the Latent Topic SimplexabstractInternational audience Kong-Aik Lee, Chang Huai You, Ville Hautamäki, Anthony Larcher, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2010 | Decoupling session variability modelling and speaker characterisationabstractThe Factor Analysis framework demonstrated its high power to model session variability during the past years. However, training the FA parameters implies to have a large amount of training data. When the size of the available database is limited, the number of components of the core statistical model, the UBM, is also limited as the UBM drives the dimension of the FA main matrix. As the size of the UBM gives directly the size of the speaker supervector (concatenation of the GMM mean parameters), it limits also the intrinsic capacity of the recognition system , reducing the performance expectation. This paper aims to withdraw this limitation by breaking the intrinsic link between the FA dimensionality and the UBM dimensionality. The session variability modelling is done on a smaller dimension compared to the UBM, which drives the discriminative power of the system. The first experimental results proposed in this paper, done using the NIST-SRE 2008 framework, are encouraging with a relative EER improvement of about 18% when a 512 components UBM is associated to a 32 components session variability modelling compared with a 32 components UBM associated with the same variability modelling. Anthony Larcher, Christophe Lévy, Driss Matrouf, Jean-François Bonastre |
INTERSPEECH | 1 |
| 2008 | Reinforced temporal structure information for embedded utterance-based speaker recognitionabstractEmbedded speaker recognition in mobile devices could involve several ergonomic constraints and a limited amount of com-puting resources. Even if they have proved their efficiency in more classical contexts, GMM/UBM based systems show their limits in such situations, with good accuracy demanding a rel-atively large quantity of speech data, but with negligible har-nessing of linguistic content. The proposed approach addresses these limitations and takes advantage from the linguistic nature of the speech material into the GMM/UBM framework by us-ing client-customised utterances. The GMM/UBM is then rein-forced with new temporal information. Experiments on the MyIdea database are performed when im-postors know the client-utterance and also when they do not, highlighting the potential of this new approach. A relative gain up to 45 % in terms of EER is achieved when impostors do not know the client utterance and performance is equivalent to the GMM/UBM baseline system in other configurations. 1. Anthony Larcher, Jean-François Bonastre, John S. D. Mason |
INTERSPEECH | 1 |
| 2008 | Short utterance-based video aided speaker recognitionabstractEmbedded speaker recognition in mobile devices could involve several ergonomic constraints and a limited amount of computing resources. Even if they have proved their efficiency in more classical contexts, GMM/UBM based systems show their limits in such situations, with good accuracy demanding a relatively large quantity of speech data, but with negligible harnessing of linguistic content. The proposed approach addresses these limitations and takes advantage of the linguistic nature of the speech material into the GMM/UBM framework by using clientcustomised utterances. Furthermore, the acoustic structure is then reinforced with video information. Experiments on the MyIdea database are performed when impostors know the client utterance and also when they do not, highlighting the potential of this new approach. A relative gain up to 47% in terms of EER is achieved when impostors do not know the client utterance and performance is equivalent to the GMM/UBM baseline system in other configurations. Anthony Larcher, Jean-François Bonastre, John S. D. Mason |
MMSP | 1 |