VLDB 2026 Research / reviewers in the wild / expert
Aghilas Sini
dblp:166/6633
· DBLP profile ↗
11ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0001-6785-6936ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scaling pseudo-labeling data for end-to-end low-resource speech translation (the case of Kurdish language)
Mohammad MohammadAmini, Aghilas Sini, Marie Tahon, Antoine Laurent |
INTERSPEECH | 2 |
| 2025 | Leveraging SSL Speech Features and Mamba for Enhanced DeepFake DetectionabstractInternational audience Hoan My Tran, Damien Lolive, David Guennec, Aghilas Sini, Arnaud Delhay, Pierre-François Marteau |
INTERSPEECH | 4 |
| 2025 | Multi-level SSL Feature Gating for Audio Deepfake DetectionabstractRecent advancements in generative AI, particularly in speech synthesis, have enabled the generation of highly natural-sounding synthetic speech that closely mimics human voices. While these innovations hold promise for applications like assistive technologies, they also pose significant risks, including misuse for fraudulent activities, identity theft, and security threats. Current research on spoofing detection countermeasures remains limited by generalization to unseen deepfake attacks and languages. To address this, we propose a gating mechanism extracting relevant feature from the speech foundation XLS-R model as a front-end feature extractor. For downstream back-end classifier, we employ Multi-kernel gated Convolution (MultiConv) to capture both local and global speech artifacts. Additionally, we introduce Centered Kernel Alignment (CKA) as a similarity metric to enforce diversity in learned features across different MultiConv layers. By integrating CKA with our gating mechanism, we hypothesize that each component helps improving the learning of distinct synthetic speech patterns. Experimental results demonstrate that our approach achieves state-of-the-art performance on in-domain benchmarks while generalizing robustly to out-of-domain datasets, including multilingual speech samples. This underscores its potential as a versatile solution for detecting evolving speech deepfake threats. Hoan My Tran, Damien Lolive, Aghilas Sini, Arnaud Delhay, Pierre-François Marteau, David Guennec |
ACM Multimedia | 3 |
| 2024 | Spoofed Speech Detection with a Focus on Speaker EmbeddingabstractInternational audience Hoan My Tran, David Guennec, Aghilas Sini, Damien Lolive, Arnaud Delhay, Pierre-François Marteau |
INTERSPEECH | 4 |
| 2022 | Investigating Inter- and Intra-speaker Voice Conversion using AudiobooksabstractAudiobook readers play with their voices to emphasize some text passages, highlight discourse changes or significant events, or in order to make listening easier and entertaining. A dialog is a central passage in audiobooks where the reader applies significant voice transformation, mainly prosodic modifications, to realize character properties and changes. However, these intra-speaker modifications are hard to reproduce with simple text-to-speech synthesis. The manner of vocalizing characters involved in a given story depends on the text style and differs from one speaker to another. In this work, this problem is investigated through the prism of voice conversion. We propose to explore modifying the narrator’s voice to fit the context of the story, such as the character who is speaking, using voice conversion. To this end, two complementary experiments are designed: the first one aims to assess the quality of our Phonetic PosteriorGrams (PPG)-based voice conversion system using parallel data. Subjective evaluations with naive raters are conducted to estimate the quality of the signal generated and the speaker similarity. The second experiment applies an intra-speaker voice conversion, considering narration passages and direct speech passages as two distinct speakers. Data are then nonparallel and the dissimilarity between character and narrator is subjectively measured. Aghilas Sini, Damien Lolive, Nelly Barbot, Pierre Alain |
LREC | 1 |
| 2022 | Phone-Level Pronunciation Scoring for L1 Using Weighted-Dynamic Time WarpingabstractThis paper presents a novel approach for phone-level pronunciation scoring. The proposed method relies on the two usual stages of pronunciation scoring: an acoustic model transcribes the spoken utterance into a phoneme sequence and then, Weighted-Dynamic Time Warping (W-DTW) is used to compare the predicted phoneme sequence against the reference one. Our approach alters the comparison process by considering Phonetic PosteriorGrams (PPG) rather than only the most probable sequence of phonemes. This led us to propose a modified W-DTW algorithm that considers the probabilities of the predicted phonemes, as well as the use of articulatory features as a proxy of phonetic similarity. The results achieved are satisfactory considering the content of the adult speech database and are comparable to well-known state-of-the-art methods. Aghilas Sini, Antoine Perquin, Damien Lolive, Arnaud Delhay |
SLT | 1 |
| 2018 | SynPaFlex-Corpus: An Expressive French Audiobooks Corpus dedicated to expressive speech synthesis
Aghilas Sini, Damien Lolive, Gaëlle Vidal, Marie Tahon, Elisabeth Delais-Roussarie |
LREC | 1 |
| 2017 | Towards confidence measures on fundamental frequency estimationsabstractThe fundamental frequency is one of the prosodic parameters, and many algorithms have been developed for estimating the fundamental frequency of speech signals. Most of them provide good results on good quality speech signals, but their performance degrades when dealing with noisy signals. Moreover, although some provide a probability for the voicing decision, none of them indicate how reliable the estimated fundamental frequency is. In this paper, we investigate the computation of a confidence (or reliability) measure on the estimated fundamental frequency values. A neural network based approach is proposed for computing the posterior probability that the estimated fundamental frequency is correct. Experiments are conducted on the PTDB-TUG pitch-tracking database, using three fundamental frequency estimation algorithms. Boyuan Deng, Denis Jouvet, Yves Laprie, Ingmar Steiner, Aghilas Sini |
ICASSP | 5 |
| 2017 | End-to-End Acoustic Feedback in Language Learning for Correcting Devoiced French Final-FricativesabstractInternational audience Sucheta Ghosh, Camille Fauth, Yves Laprie, Aghilas Sini |
INTERSPEECH | 4 |
| 2016 | L1-L2 Interference: The Case of Final Devoicing of French Voiced Fricatives in Final Position by German LearnersabstractInternational audience Sucheta Ghosh, Camille Fauth, Aghilas Sini, Yves Laprie |
INTERSPEECH | 3 |
| 2015 | Audio source localization by optimal control of a mobile robotabstractWe consider the task of audio source localization using a microphone array on a mobile robot. Active localization algorithms have been proposed in the literature that can estimate the 3D position of a source by fusing the measurements taken for different poses of the robot. The robot movements are typically fixed, however, or they obey heuristic strategies, such as turning the head and moving towards the source, which may be suboptimal. In this paper, we propose to control the robot movements so as to locate the source as quickly as possible. We represent the belief about the source position by a discrete grid and we introduce a dynamic programming algorithm to find the optimal robot motion minimizing the entropy of the grid. We report initial results in a real environment. Emmanuel Vincent 0001, Aghilas Sini, François Charpillet |
ICASSP | 2 |