VLDB 2026 Research / reviewers in the wild / expert
Réda Dehak
dblp:84/8053 · also Sidi Mohammed Réda Dehak
· DBLP profile ↗
21ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0002-4078-7261ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 5 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Data Selection Effects on Self-Supervised Learning of Audio Representations for French Audiovisual Broadcasts
Valentin Pelloin, Lina Bekkali, Réda Dehak, David Doukhan |
LREC | 3 |
| 2026 | Self-Supervised Learning for Speaker Recognition: A study and review
Théo Lepage, Réda Dehak |
Speech Commun. | 2 |
| 2025 | SSPS: Self-Supervised Positive Sampling for Robust Self-Supervised Speaker VerificationabstractInternational audience Théo Lepage, Réda Dehak |
INTERSPEECH | 2 |
| 2024 | InaGVAD : A Challenging French TV and Radio Corpus Annotated for Speech Activity Detection and Speaker Gender SegmentationabstractInaGVAD is an audio corpus collected from 10 French radio and 18 TV channels categorized into 4 groups: generalist radio, music radio, news TV, and generalist TV. It contains 277 1-minute-long annotated recordings aimed at representing the acoustic diversity of French audiovisual programs and was primarily designed to build systems able to monitor men’s and women’s speaking time in media. inaGVAD is provided with Voice Activity Detection (VAD) and Speaker Gender Segmentation (SGS) annotations extended with overlap, speaker traits (gender, age, voice quality), and 10 non-speech event categories. Annotation distributions are detailed for each channel category. This dataset is partitioned into a 1h development and a 3h37 test subset, allowing fair and reproducible system evaluation. A benchmark of 6 freely available VAD software is presented, showing diverse abilities based on channel and non-speech event categories. Two existing SGS systems are evaluated on the corpus and compared against a baseline X-vector transfer learning strategy, trained on the development subset. Results demonstrate that our proposal, trained on a single - but diverse - hour of data, achieved competitive SGS results. The entire inaGVAD package; including corpus, annotations, evaluation scripts, and baseline training code; is made freely accessible, fostering future advancement in the domain. David Doukhan, Christine Maertens, William Le Personnic, Ludovic Speroni, Réda Dehak |
LREC/COLING | 5 |
| 2024 | Towards Supervised Performance on Speaker Verification with Self-Supervised Learning by Leveraging Large-Scale ASR ModelsabstractInternational audience Victor Miara, Théo Lepage, Réda Dehak |
INTERSPEECH | 3 |
| 2023 | Experimenting with Additive Margins for Contrastive Self-Supervised Speaker VerificationabstractInternational audience Théo Lepage, Réda Dehak |
INTERSPEECH | 2 |
| 2022 | Label-Efficient Self-Supervised Speaker Verification With Information Maximization and Contrastive LearningabstractInternational audience Théo Lepage, Réda Dehak |
INTERSPEECH | 2 |
| 2020 | State-of-the-art speaker recognition with neural network embeddings in NIST SRE18 and Speakers in the Wild evaluations
Jesús Villalba 0001, Nanxin Chen, David Snyder, Daniel Garcia-Romero, Alan McCree, Gregory Sell, Jonas Borgstrom, L. Paola García-Perera, Fred Richardson, Réda Dehak, Pedro A. Torres-Carrasquillo, Najim Dehak |
Comput. Speech Lang. | 10 |
| 2019 | State-of-the-Art Speaker Recognition for Telephone and Video Speech: The JHU-MIT Submission for NIST SRE18
Jesús Villalba 0001, Nanxin Chen, David Snyder, Daniel Garcia-Romero, Alan McCree, Gregory Sell, Jonas Borgstrom, Fred Richardson, Suwon Shon, François Grondin, Réda Dehak, L. Paola García-Perera, Daniel Povey, Pedro A. Torres-Carrasquillo, Sanjeev Khudanpur, Najim Dehak |
INTERSPEECH | 11 |
| 2017 | The MIT-LL, JHU and LRDE NIST 2016 Speaker Recognition Evaluation System
Pedro A. Torres-Carrasquillo, Fred Richardson, Shahan C. Nercessian, Douglas E. Sturim, William M. Campbell, Youngjune Gwon, Swaroop Vattam, Najim Dehak, Sri Harish Reddy Mallidi, Phani S. Nidadavolu, Réda Dehak |
INTERSPEECH | 12 |
| 2013 | Unsupervised Methods for Speaker Diarization: An Integrated and Iterative ApproachabstractIn speaker diarization, standard approaches typically perform speaker clustering on some initial segmentation before refining the segment boundaries in a re-segmentation step to obtain a final diarization hypothesis. In this paper, we integrate an improved clustering method with an existing re-segmentation algorithm and, in iterative fashion, optimize both speaker cluster assignments and segmentation boundaries jointly. For clustering, we extend our previous research using factor analysis for speaker modeling. In continuing to take advantage of the effectiveness of factor analysis as a front-end for extracting speaker-specific features (i.e., i-vectors), we develop a probabilistic approach to speaker clustering by applying a Bayesian Gaussian Mixture Model (GMM) to principal component analysis (PCA)-processed i-vectors. We then utilize information at different temporal resolutions to arrive at an iterative optimization scheme that, in alternating between clustering and re-segmentation steps, demonstrates the ability to improve both speaker cluster assignments and segmentation boundaries in an unsupervised manner. Our proposed methods attain results that are comparable to those of a state-of-the-art benchmark set on the multi-speaker CallHome telephone corpus. We further compare our system with a Bayesian nonparametric approach to diarization and attempt to reconcile their differences in both methodology and performance. Stephen H. Shum, Najim Dehak, Réda Dehak, James R. Glass |
IEEE Trans. Speech Audio Process. | 3 |
| 2011 | A channel-blind system for speaker verificationabstractThe majority of speaker verification systems proposed in the NIST speaker recognition evaluation are conditioned on the type of data to be processed: telephone or microphone. In this paper, we propose a new speaker verification system that can be applied to both types of data. This system, named blind system, is based on an extension of the total variability framework. Recognition results with the pro posed channel-independent system are comparable to state of the art systems that require conditioning on the channel type. Another ad vantage of our proposed system is that it allows for combining data from multiple channels in the same visualization in order to explore the effects of different microphones and collection environments. Najim Dehak, Zahi N. Karam, Douglas A. Reynolds, Réda Dehak, William M. Campbell, James R. Glass |
ICASSP | 4 |
| 2011 | Language Recognition via i-vectors and Dimensionality ReductionabstractIn this paper, a new language identification system is presented based on the total variability approach previously developed in the field of speaker identification. Various techniques are em-ployed to extract the most salient features in the lower dimen-sional i-vector space and the system developed results in excel-lent performance on the 2009 LRE evaluation set without the need for any post-processing or backend techniques. Additional performance gains are observed when the system is combined with other acoustic systems. Najim Dehak, Pedro A. Torres-Carrasquillo, Douglas A. Reynolds, Réda Dehak |
INTERSPEECH | 4 |
| 2011 | Front-End Factor Analysis for Speaker VerificationabstractThis paper presents an extension of our previous work which proposes a new speaker representation for speaker verification. In this modeling, a new low-dimensional speaker- and channel-dependent space is defined using a simple factor analysis. This space is named the total variability space because it models both speaker and channel variabilities. Two speaker verification systems are proposed which use this new representation. The first system is a support vector machine-based system that uses the cosine kernel to estimate the similarity between the input data. The second system directly uses the cosine similarity as the final decision score. We tested three channel compensation techniques in the total variability space, which are within-class covariance normalization (WCCN), linear discriminate analysis (LDA), and nuisance attribute projection (NAP). We found that the best results are obtained when LDA is followed by WCCN. We achieved an equal error rate (EER) of 1.12% and MinDCF of 0.0094 using the cosine distance scoring on the male English trials of the core condition of the NIST 2008 Speaker Recognition Evaluation dataset. We also obtained 4% absolute EER improvement for both-gender trials on the 10 s-10 s condition compared to the classical joint factor analysis scoring. Najim Dehak, Patrick Kenny, Réda Dehak, Pierre Dumouchel, Pierre Ouellet |
IEEE Trans. Speech Audio Process. | 3 |
| 2009 | Support vector machines and Joint Factor Analysis for speaker verificationabstractThis article presents several techniques to combine between support vector machines (SVM) and joint factor analysis (JFA) model for speaker verification. In this combination, the SVMs are applied to different sources of information produced by the JFA. These informations are the Gaussian mixture model supervectors and speakers and common factors. We found that using SVM in JFA factors gave the best results especially when within class covariance normalization method is applied in order to compensate for the channel effect. The new combination results are comparable to other classical JFA scoring techniques. Najim Dehak, Patrick Kenny, Réda Dehak, Ondrej Glembek, Pierre Dumouchel, Lukás Burget, Valiantsina Hubeika, Fabio Castaldo |
ICASSP | 3 |
| 2009 | Support vector machines versus fast scoring in the low-dimensional total variability space for speaker verificationabstractThis paper presents a new speaker verification system architecture based on Joint Factor Analysis (JFA) as feature extractor. In this modeling, the JFA is used to define a new low-dimensional space named the total variability factor space, instead of both channel and speaker variability spaces for the classical JFA. The main contribution in this approach, is the use of the cosine kernel in the new total factor space to design two different systems: the first system is Support Vector Machines based, and the second one uses directly this kernel as a decision score. This last scoring method makes the process faster and less computation complex compared to others classical methods. We tested several intersession compensation methods in total factors, and we found that the combination of Linear Discriminate Analysis and Within Class Covariance Normalization achieved the best performance. We achieved a remarkable results using fast scoring method based only on cosine kernel especially for male trials, we yield an EER of 1.12% and MinDCF of 0.0094 on the English trials of the NIST 2008 SRE dataset. Index Terms: Total variability space, cosine kernel, fast scoring, support vector machines. Najim Dehak, Réda Dehak, Patrick Kenny, Niko Brümmer, Pierre Ouellet, Pierre Dumouchel |
INTERSPEECH | 2 |
| 2009 | Cepstral and long-term features for emotion recognitionabstractIn this paper, we describe systems that were developed for the Open Performance Sub-Challenge of the INTERSPEECH 2009 Emotion Challenge. We participate in both two-class and fiveclass emotion detection. For the two-class problem, the best performance is obtained by logistic regression fusion of three systems. These systems use short- and long-term speech features. Fusion allowed to an absolute improvement of 2:6% on the unweighted recall value compared with [1]. For the fiveclass problem, we submitted two individual systems: cepstral GMM vs. long-term GMM-UBM. The best result comes from a cepstral GMM and produces an absolute improvement of 3:5% compared to [6]. Pierre Dumouchel, Najim Dehak, Yazid Attabi, Réda Dehak, Narjès Boufaden |
INTERSPEECH | 4 |
| 2007 | Linear and non linear kernel GMM supervector machines for speaker verificationabstractThis paper presents a comparison between Support Vector Machines (SVM) speaker verification systems based on linear and non linear kernels defined in GMM supervector space. We describe how these kernel functions are related and we show how the nuisance attribute projection (NAP) technique can be used with both of these kernels to deal with the session variability problem. We demonstrate the importance of GMM model normalization (M-Norm) especially for the non linear kernel. All our experiments were performed on the core condition of NIST 2006 speaker recognition evaluation (all trials). Our best results (an equal error rate of 6.3%) were obtained using NAP and GMM model normalization with the non linear kernel. Réda Dehak, Najim Dehak, Patrick Kenny, Pierre Dumouchel |
INTERSPEECH | 1 |
| 2005 | Spatial Reasoning with Incomplete Information on Relative PositioningabstractThis paper describes a probabilistic method of inferring the position of a point with respect to a reference point knowing their relative spatial position to a third point. We address this problem in the case of incomplete information where only the angular spatial relationships are known. The use of probabilistic representations allows us to model prior knowledge. We derive exact formulae expressing the conditional probability of the position given the two known angles, in typical cases: uniform or Gaussian random prior distributions within rectangular or circular regions. This result is illustrated with respect to two different simulations: The first is devoted to the localization of a mobile phone using only angular relationships, the second, to geopositioning within a city. This last example uses angular relationships and some additional knowledge about the position. Réda Dehak, Isabelle Bloch, Henri Maître |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2004 | A novel method to fight the non-line-of-sight error in AOA measurements for mobile locationabstractIn this contribution, a mobile location method is introduced using angle of arrival's (AOAs) in UMTS-FDD systems with measurements from at least two different base-stations. Although computationally efficient, this method enhances state of the art algorithms based on a simple trilateration and takes into account error measurements caused by non-line-of-sight (NLOS) and near-far effect. The new method attributes an index of confidence for each measure, in order to allow the mobile to select the two most reliable measures and not to use all measures, equally. Emmanuèle Grosicki, Karim Abed-Meraim, Réda Dehak |
ICC | 3 |
| 2001 | Inference of directional spatial relationship between points: a probabilistic approachabstractThis paper develops an evaluation of the position probability of a point C which is known to be in a direction /spl beta/ with respect to a point B, itself in the direction /spl alpha/ with respect to another point A. The obtained results can be used in the problem of inference of directional relationships in the case of spatial reasoning. Réda Dehak, Isabelle Bloch, Henri Maître |
ICIP (3) | 1 |