Elie Khoury 0001

dblp:05/1112-1 · also Elie el Khoury 0001 · DBLP profile ↗
← Back
31ranked-venue papers
9as first author
11since 2021 · last 2025
0000-0001-9568-3729ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 29 · 8 first-author · 11 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 3 since 2021Security and privacy · 2Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Investigating voiced and unvoiced regions of speech for audio deepfake detection
abstract
Deep neural network based deepfake detection systems have achieved high levels of accuracy on benchmark datasets and competitions. However, most models lack interpretability. It is challenging to extract reasoning from the network that can convince the human evaluator to trust the decision. Humans often rely on acoustic cues like unnatural pitch jitter, robotic intonation, acoustic artifacts, and unnatural sounding fricatives to judge the quality of the synthetic audio. This study explores the role played by the voiced and unvoiced regions of speech in discriminating synthetic from bonafide speech. A measure of signal periodicity is used to analyze speech into voiced and unvoiced components. Then, the graph attention based AASIST detection system is trained independently on each component. This work compares the accuracy of deepfake detection system using voiced and unvoiced components and analyzes the results on the MLAAD dataset. Our results show that unvoiced regions are particularly more effective in distinguishing synthetic (deepfake) speech from bonafide, and achieves an equal error rate of 6.62%. When combined with voice regions through score-level fusion, the overall performance improves further, yielding a 5.82% EER, a relative improvement of 49% over the baseline system that uses the full audio.
Ganesh Sivaraman, Hemlata Tak, Elie Khoury 0001
ICASSP3
2025 Open-Set Source Tracing of Audio Deepfake Systems
Nicholas Klein, Hemlata Tak, Elie Khoury 0001
INTERSPEECH3
2025 Pindrop it! Audio and Visual Deepfake Countermeasures for Robust Detection and Fine-Grained Localization
abstract
The field of visual and audio generation is burgeoning with new state-of-the-art methods. This rapid proliferation of new techniques underscores the need for robust solutions for detecting synthetic content in videos. In particular, when fine-grained alterations via localized manipulations are performed in visual, audio, or both domains, these subtle modifications add challenges to the detection algorithms. This paper presents solutions for the problems of deepfake video classification and localization. The methods were submitted to the ACM 1M Deepfakes Detection Challenge, achieving the best performance in the temporal localization task and a top four ranking in the classification task for the TestA split of the evaluation dataset.
Nicholas Klein, Hemlata Tak, James Fullwood, Krishna Regmi, Leonidas Spinoulas, Ganesh Sivaraman, Elie Khoury 0001
ACM Multimedia8
2024 Source Tracing of Audio Deepfake Systems
Nicholas Klein, Hemlata Tak, Ricardo Casal, Elie Khoury 0001
INTERSPEECH5
2022 Speaker Embedding Conversion for Backward and Cross-Channel Compatibility
abstract
The accuracy of automatic speaker verification (ASV) systems has shown tremendous improvements due to the recent breakthroughs in low-rank speaker representations and deep learning techniques, leading to the success of ASV in real-world applications from call centers to mobile applications and smart devices. Particularly, some ASV providers have been migrating their legacy systems from the traditional GMM based i-vector paradigm to the deep learning based x-vector paradigm. Additionally, some of them are in need of implementing simultaneously different systems for different use cases such as 8 kHz over the phone channel and 16 kHz on virtual assistants. In either cases, the speaker embeddings extracted from one ASV system are often not compatible with another ASV system. This makes the process of interchangeability between systems very cumbersome and costly. In this paper, we address this issue by proposing a highly efficient speaker embedding converter that transforms a speaker embedding extracted from system A into a speaker embedding that can be used by system B. We evaluate the performance of the embedding converter for i-vector to x-vector upgrade scenario and for cross channel compatibility scenario. In both scenarios, we show that the proposed system achieves very low and compelling equal error rates.
Elie Khoury 0001
ICASSP2
2022 Distribution Learning for Age Estimation from Speech
abstract
Age estimation from speech is becoming important with increasing usage of the voice channel. Call centers can use age estimates to influence call routing or to provide security by comparison with the speaker's age on-file. Voice assistants can use it for parental control applications. The problem of age estimation from speech has been often viewed as a regression or classification problem. However these methods do not explicitly incorporate ordinal ranking or uncertainty in age estimation that humans often do. In this work, we hypothesize that the age follows a normal distribution centered around the real age with a particular confidence interval. We investigate three different distribution learning losses, namely KL divergence, GJM distance and mean-and-variance loss. Cross-dataset experiments were conducted on the NIST SRE08/10 and AgeVoxCeleb data, and their results show that the distribution learning methods are very competitive and in most cases better than traditional approaches.
Amruta Saraf, Elie Khoury 0001
ICASSP2
2022 Unsupervised Model Adaptation for End-to-End ASR
abstract
End-to-end (E2E) Automatic Speech Recognition (ASR) systems are widely applied in various devices and communication domains. However, state-of-the-art ASR systems are known to underperform when there is a mismatch in the training and test domains. As a result, acoustic models deployed in production are often adapted to the target domain to improve accuracy. This paper proposes a method to perform unsupervised model adaptation for E2E ASR using first-pass transcriptions of adaptation data produced by the baseline ASR model itself. The paper proposes two transcription confidence measures that can be used to select an optimal in-domain adaptation set. Experiments were performed using the Quartznet ASR architecture on the HarperValleyBank corpus. Results show that the unsupervised adaptation technique with the confidence measure based data selection results in a 8% absolute reduction in word error rate on the HarperValleyBank test set. The proposed method can be applied to any E2E ASR system and is suitable for model adaptation on call center audio with little to no manual transcription.
Ganesh Sivaraman, Ricardo Casal, Matt Garland, Elie Khoury 0001
ICASSP4
2022 Confidence Measure for Automatic Age Estimation From Speech
Amruta Saraf, Ganesh Sivaraman, Elie Khoury 0001
INTERSPEECH3
2022 A Zero-Shot Approach to Identifying Children's Speech in Automatic Gender Classification
abstract
Detecting whether a speech utterance belongs to an adult male, adult female or a child category, also known as male-female-child (MFC) classification is particularly challenging due to two main reasons - paucity of children's speech data, and high variability in children's speech due to developmental changes. It is difficult to obtain speech datasets with children's voices due to privacy reasons. This paper explores a zero-shot learning approach to MFC classification. Different algorithms are explored to create artificial childlike voices from adult voices. Methods such as pitch shifting, Vocal Tract Length Perturbation, and Segmental Warping Perturbation are used to create synthetic childlike speech for the MFC classification task. Speaker embeddings extracted from a DNN based speaker recognition system are used as features for MFC classification. Compared to a pitch frequency based baseline MFC classifier, the proposed method improves the child classification accuracy by 47%.
Amruta Saraf, Ganesh Sivaraman, Elie Khoury 0001
SLT3
2021 Spoofprint: A New Paradigm for Spoofing Attacks Detection
abstract
With the development of voice spoofing techniques, voice spoofing attacks have become one of the main threats to automatic speaker verification (ASV) systems. Traditionally, researchers tend to treat this problem as a binary classification task. A binary classifier is typically trained using machine learning (including deep learning) algorithms to determine whether a given audio clip is bonafide or spoofed. This approach is effective on detecting spoofing attacks that are generated by known voice spoofing techniques. However, in practical scenarios, new types of spoofing technologies are emerging rapidly. It is impossible to include all types of spoofing technologies into the training dataset, and thus it is desired that the detection system can generalize to unseen spoofing techniques. In this paper, we propose a new paradigm for spoofing attacks detection called Spoofprint. Instead of using a binary classifier to detect spoofed audio, Spoofprint uses a paradigm similar to ASV systems and involves an enrollment phase and a verification phase. We evaluate the performance on the original and noisy versions of ASVspoof 2019 logical access (LA) dataset. The results show that the proposed Spoofprint paradigm is effective on detecting unknown type of attacks and is often superior to the latest state-of-the-art.
Elie Khoury 0001
SLT2
2021 Improving Speaker Recognition with Quality Indicators
abstract
Nuisance factors such as short duration, noise and transmission conditions still pose accuracy challenges to state-of-the-art automatic speaker verification (ASV) systems. To address this problem, we propose a no reference system that consumes quality indicators encapsulating information about duration of speech, acoustic events and codec artifacts. These quality indicators are used as estimates to measure how close a given speech utterance would be to a high-quality speech segment uttered by the same speaker. The proposed measures when fused with a baseline ASV system are found to improve the performance of speaker recognition. The experimental study carried on a modified version of the NIST SRE 2019 dataset shows a relative decrease of 9.6% in equal error rate (EER) compared to the baseline.
Hrishikesh Rao 0001, Kedar Phatak, Elie Khoury 0001
SLT3
2019 Pindrop Labs' Submission to the First Multi-Target Speaker Detection and Identification Challenge
Elie Khoury 0001, Khaled Lakhdhar, Andrew Vaughan, Ganesh Sivaraman, Parav Nagarsheth
INTERSPEECH1
2018 Speech Synthesis in the Wild
Ganesh Sivaraman, Parav Nagarsheth, Elie Khoury 0001
INTERSPEECH3
2017 Replay Attack Detection Using DNN for Channel Discrimination
Parav Nagarsheth, Elie Khoury 0001, Kailash Patil, Matt Garland
INTERSPEECH2
2016 HAPPY Team Entry to NIST OpenSAD Challenge: A Fusion of Short-Term Unsupervised and Segment i-Vector Based Speech Activity Detectors
abstract
Speech activity detection (SAD), the task of locating speech segments from a given recording, remains challenging under acoustically degraded conditions. In 2015, National Institute of Standards and Technology (NIST) coordinated OpenSAD bench-mark. We summarize “HAPPY” team effort to OpenSAD. SADs come in both unsupervised and supervised flavors, the latter requiring a labeled training set. Our solution fuses six base SADs (2 supervised and 4 unsupervised). The individually best SAD, in terms of detection cost function (DCF), is supervised and uses adaptive segmentation with i-vectors to represent the segments. Fusion of the six base SADs yields a relative decrease of 9.3% in DCF over this SAD. Further, relative decrease of 17.4% is obtained by incorporating channel detection side information.
Tomi Kinnunen, Alexey Sholokhov, Elie Khoury 0001, Dennis Alexander Lehmann Thomsen, Md. Sahidullah, Zheng-Hua Tan
INTERSPEECH3
2015 Joint Speaker Verification and Antispoofing in the i-Vector Space
abstract
Any biometric recognizer is vulnerable to spoofing attacks and hence voice biometric, also called automatic speaker verification (ASV), is no exception; replay, synthesis, and conversion attacks all provoke false acceptances unless countermeasures are used. We focus on voice conversion (VC) attacks considered as one of the most challenging for modern recognition systems. To detect spoofing, most existing countermeasures assume explicit or implicit knowledge of a particular VC system and focus on designing discriminative features. In this paper, we explore back-end generative models for more generalized countermeasures. In particular, we model synthesis-channel subspace to perform speaker verification and antispoofing jointly in the i-vector space, which is a well-established technique for speaker modeling. It enables us to integrate speaker verification and antispoofing tasks into one system without any fusion techniques. To validate the proposed approach, we study vocoder-matched and vocoder-mismatched ASV and VC spoofing detection on the NIST 2006 speaker recognition evaluation data set. Promising results are obtained for standalone countermeasures as well as their combination with ASV systems using score fusion and joint approach.
Aleksandr Sizov, Elie Khoury 0001, Tomi Kinnunen, Zhizheng Wu 0001, Sébastien Marcel
IEEE Trans. Inf. Forensics Secur.2
2014 A conditional random field approach for audio-visual people diarization
abstract
We investigate the problem of audio-visual (AV) person diarization in broadcast data. That is, automatically associate the faces and voices of people and determine when they appear or speak in the video. The contributions are twofolds. First, we formulate the problem within a novel CRF framework that simultaneously performs the AV association of voices and face clusters to build AV person models, and the joint segmentation of the audio and visual streams using a set of AV cues and their association strength. Secondly, we use for this AV association strength a score that does not only rely on lips activity, but also on contextual visual information (face size, position, number of detected faces,...) that leads to more reliable association measures. Experiments on 6 hours of broadcast data show that our framework is able to improve the AV-person diarization especially for speaker segments erroneously labeled in the mono-modal case.
Paul Gay, Elie Khoury 0001, Sylvain Meignier, Jean-Marc Odobez, Paul Deléglise
ICASSP2
2014 Spear: An open source toolbox for speaker recognition based on Bob
abstract
In this paper, we introduce Spear, an open source and extensible toolbox for state-of-the-art speaker recognition. This toolbox is built on top of Bob, a free signal processing and machine learning library. Spear implements a set of complete speaker recognition toolchains, including all the processing stages from the front-end feature extractor to the final steps of decision and evaluation. Several state-of-the-art modeling techniques are included, such as Gaussian mixture models, inter-session variability, joint factor analysis and total variability (i-vectors). Furthermore, the toolchains can be easily evaluated on well-known databases such as NIST SRE and MOBIO. As a proof of concept, an experimental comparison of different modeling techniques is conducted on the MOBIO database.
Elie Khoury 0001, Laurent El Shafey, Sébastien Marcel
ICASSP1
2014 Audio-visual gender recognition in uncontrolled environment using variability modeling techniques
abstract
The problem of gender recognition using visual and acoustic cues has recently received significant attention. This paper explores the use of Total Variability (i-vectors) and Inter-Session Variability (ISV) modeling techniques for both unimodal and bimodal gender recognition, and compares them to several state-of-the-art algorithms. The experimental evaluation is conducted on the FERET and LFW databases for face-based gender recognition, on the NIST-SRE database for audio-based gender recognition, and on the MOBIO database for audio-visual gender recognition. Results on LFW show that the i-vectors technique outperforms state-of-the-art algorithms, which are based on Support Vector Machines (SVM) applied either on raw pixels, on Local Binary Patterns (LBP) or on Gabor filters, with an accuracy rate of about 95%. Results on NIST-SRE show that the i-vectors system is also superior to state-of-the-art GMM-based gender recognition systems, with a relative gain of about 11%. Finally, results on MOBIO show that i-vectors and ISV also take advantage of combining visual and acoustic cues using logistic regression. The resulting bimodal systems achieve accuracy rates of about 98%.
Laurent El Shafey, Elie Khoury 0001, Sébastien Marcel
IJCB2
2014 A conditional random field approach for face identification in broadcast news using overlaid text
abstract
We investigate the problem of face identification in broadcast programs where people names are obtained from text overlays automatically processed with Optical Character Recognition (OCR) and further linked to the faces throughout the video. To solve the face-name association and propagation, we propose a novel approach that combines the positive effects of two Conditional Random Field (CRF) models: a CRF for person diarization (joint temporal segmentation and association of voices and faces) that benefit from the combination of multiple cues including as main contributions the use of identification sources (OCR appearances) and recurrent local face visual background (LFB) playing the role of a namedness feature; a second CRF for the joint identification of the person clusters that improves identification performance thanks to the use of further diarization statistics. Experiments conducted on a recent and substantial public dataset of 7 different shows demonstrate the interest and complementarity of the different modeling steps and information sources, leading to state of the art results.
Paul Gay, Elie Khoury 0001, Sylvain Meignier, Jean-Marc Odobez, Paul Deléglise
ICIP2
2014 Dialect levelling in Finnish: a universal speech attribute approach
abstract
We adopt automatic language recognition methods to study dialect levelling - a phenomenon that leads to reduced structural differences among dialects in a given spoken language. In terms of dialect characterisation, levelling is a nuisance variable that adversely affects recognition accuracy: The more similar two dialects are, the harder it is to set them apart. We address levelling in Finnish regional dialects using a new SAPU (Satakunta in Speech) corpus containing material from Satakunta (South-Western Finland) between 2007 and 2013. To define a compact and universal set of sound units to characterize dialects, we adopt speech attributes features, namely manner and place of articulation. It will be shown that speech attribute distributions can indeed characterise differences among dialects. Experiments with an i-vector system suggest that (1) the attribute features achieve higher dialect recognition accuracy and (2) they are less sensitive against age-related levelling in comparison to traditional spectral approach.
Hamid Behravan, Ville Hautamäki, Sabato Marco Siniscalchi, Elie Khoury 0001, Tommi Kurki, Tomi Kinnunen, Chin-Hui Lee 0001
INTERSPEECH4
2014 Introducing i-vectors for joint anti-spoofing and speaker verification
abstract
Any biometric recognizer is vulnerable to direct spoofing attacks and automatic speaker verification (ASV) is no exception; replay, synthesis and conversion attacks all provoke false acceptances unless countermeasures are used.We focus on voice conversion (VC) attacks.Most existing countermeasures use full knowledge of a particular VC system to detect spoofing.We study a potentially more universal approach involving generative modeling perspective.Specifically, we adopt standard ivector representation and probabilistic linear discriminant analysis (PLDA) back-end for joint operation of spoofing attack detector and ASV system.As a proof of concept, we study a vocoder-mismatched ASV and VC attack detection approach on the NIST 2006 speaker recognition evaluation corpus.We report stand-alone accuracy of both the ASV and countermeasure systems as well as their combination using score fusion and joint approach.The method holds promise.
Elie Khoury 0001, Tomi Kinnunen, Aleksandr Sizov, Zhizheng Wu 0001, Sébastien Marcel
INTERSPEECH1
2014 Bi-modal biometric authentication on mobile phones in challenging conditions
Elie Khoury 0001, Laurent El Shafey, Chris McCool, Manuel Günther, Sébastien Marcel
Image Vis. Comput.1
2014 Audiovisual diarization of people in video content
Elie Khoury 0001, Christine Sénac, Philippe Joly
Multim. Tools Appl.1
2013 An open-source state-of-the-art toolbox for broadcast news diarization
abstract
International audience
Mickael Rouvier, Grégor Dupuy, Paul Gay, Elie Khoury 0001, Téva Merlin, Sylvain Meignier
INTERSPEECH4
2013 I4u submission to NIST SRE 2012: a large-scale collaborative effort for noise-robust speaker verification
abstract
I4U is a joint entry of nine research Institutes and Universities across 4 continents to NIST SRE 2012. It started with a brief discussion during the Odyssey 2012 workshop in Singapore. An online discussion group was soon set up, providing a discussion platform for different issues surrounding NIST SRE’12. Noisy test segments, uneven multi-session training, variable enrollment duration, and the issue of open-set identification were actively discussed leading to various solutions integrated to the I4U submission. The joint submission and several of its 17 sub-systems were among top-performing systems. We summarize the lessons learnt from this large-scale effort.
Rahim Saeidi, Kong-Aik Lee, Tomi Kinnunen, Tawfik Hasan, Benoit G. B. Fauve, Pierre-Michel Bousquet, Elie Khoury 0001, Pablo Luis Sordo Martinez, Jia Min Karen Kua, Chang Huai You, Hanwu Sun, Anthony Larcher, Padmanabhan Rajan, Ville Hautamäki, Cemal Hanilçi, Billy Braithwaite, Rosa González Hautamäki, Seyed Omid Sadjadi, Gang Liu 0001, Hynek Boril, Navid Shokouhi, Driss Matrouf, Laurent El Shafey, Pejman Mowlaee, Julien Epps, Tharmarajah Thiruvaran, David A. van Leeuwen, Bin Ma 0001, Haizhou Li 0001, John H. L. Hansen, Jean-François Bonastre, Sébastien Marcel, John S. D. Mason, Eliathamby Ambikairajah
INTERSPEECH7
2013 Fusing matching and biometric similarity measures for face diarization in video
abstract
This paper addresses face diarization in videos, that is, deciding which face appears and when in the video. To achieve this face-track clustering task, we propose a hierarchical approach combining the strength of two complementary measures: (i) a pairwise matching similarity relying on local interest points allowing the accurate clustering of faces tracks captured in similar conditions, a situation typically found in temporally close shots of broadcast videos or in talk-shows; (ii) a biometric cross-likelihood ratio similarity measure relying on Gaussian Mixture Models (GMMs) modeling the distribution of densely sampled local features (Discrete Cosine Transform (DCT) coefficients), that better handle appearance variability. Experiments carried out on a public video dataset and on the data from the French REPERE challenge demonstrate the effectiveness of our approach in comparison with state-of-the-art methods.
Elie Khoury 0001, Paul Gay, Jean-Marc Odobez
ICMR1
2012 Combining transcription-based and acoustic-based speaker identifications for broadcast news
abstract
In this paper, we consider the issue of speaker identification within audio records of broadcast news. The speaker identity information is extracted from both transcript-based and acoustic-based speaker identification systems. This information is combined in the belief functions framework, which makes coherent the knowledge representation of the problem. The Kuhn-Munkres algorithm is used to optimize the assignment problem of speaker identities and speaker clusters. Experiments carried out on French broadcast news from the French evaluation campaign ESTER show the efficiency of the proposed combination method.
Elie Khoury 0001, Antoine Laurent, Sylvain Meignier, Simon Petit-Renaud
ICASSP1
2010 The IMMED project: wearable video monitoring of people with age dementia
abstract
In this paper, we describe a new application for multimedia indexing, using a system that monitors the instrumental activities of daily living to assess the cognitive decline caused by dementia. The system is composed of a wearable camera device designed to capture audio and video data of the instrumental activities of a patient, which is leveraged with multimedia indexing techniques in order to allow medical specialists to analyze several hour long observation shots efficiently.
Rémi Mégret, Vladislavs Dovgalecs, Hazem Wannous, Svebor Karaman, Jenny Benois-Pineau, Elie Khoury 0001, Julien Pinquier, Philippe Joly, Régine André-Obrecht, Yann Gaëstel, Jean-François Dartigues
ACM Multimedia6
2009 Improved speaker diarization system for meetings
abstract
In this paper, we investigate new approaches to improve speech activity detection, speaker segmentation and speaker clustering. The main idea behind them is to deal with the problem of speaker diarization for meetings where error rates are relatively high. In opposition to existing methods, a new iterative scheme is proposed considering those three tasks as only one problem. New bidirectional source segmentation is proposed based on the GLR/BIC method. The well-known BIC clustering is also reviewed and a new unsupervised post-processing is added to increase clusters purity. Those new proposals applied on meeting data show a relative improvement of about 40% compared to a standard speaker diarization system.
Elie Khoury 0001, Christine Sénac, Julien Pinquier
ICASSP1
2007 Speaker Diarization: Towards a More Robust and Portable System
abstract
In this paper, we describe a new method for speaker segmentation and clustering of an audio document. For the segmentation phase, we combine the generalized likelihood ratio (GLR) and the Bayesian information criterion (BIC) in a way that avoids most of the parameters tuning. For the clustering phase, we use an existing approach that utilizes the eigen vector space model (EVSM) with a bottom-up hierarchical grouping but we make some improvements by introducing prosodic information. Evaluation is done on the audio database of the ESTER evaluation campaign for the rich transcription of French Broadcast news. Results show that our method which operates without any a priori knowledge about speakers is suitable for speaker diarization as it outperforms the traditional ones with an overall diarization error rate (DER) of 16.72%.
Elie Khoury 0001, Christine Sénac, Régine André-Obrecht
ICASSP (4)1