VLDB 2026 Research / reviewers in the wild / expert
Arnaud Delhay
dblp:54/3670
· DBLP profile ↗
16ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0001-6795-7999ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Paraphrase Generation Evaluation Powered by an LLM: A Semantic Metric, Not a Lexical OneabstractEvaluating automatic paraphrase production systems is a difficult task as it involves, among other things, assessing the semantic proximity between two sentences. Usual measures are based on lexical distances, or at least on semantic embedding alignments. The rise of Large Language Models (LLM) has provided tools to model relationships within a text thanks to the attention mechanism. In this article, we introduce ParaPLUIE, a new measure based on a log likelihood ratio from an LLM, to assess the quality of a potential paraphrase. This measure is compared with usual measures on two known by the NLP community datasets prior to this study. Three new small datasets have been built to allow metrics to be compared in different scenario and to avoid data contamination bias. According to evaluations, the proposed measure is better for sorting pairs of sentences by semantic proximity. In particular, it is much more independent to lexical distance and provides an interpretable classification threshold between paraphrases and non-paraphrases. Quentin Lemesle, Jonathan Chevelu, Damien Lolive, Arnaud Delhay, Nelly Barbot |
COLING | 5 |
| 2025 | Audio Deepfake Source Tracing using Multi-Attribute Open-Set Identification and VerificationabstractInternational audience Pierre Falez, Tony Marteau, Damien Lolive, Arnaud Delhay |
INTERSPEECH | 4 |
| 2025 | Leveraging SSL Speech Features and Mamba for Enhanced DeepFake DetectionabstractInternational audience Hoan My Tran, Damien Lolive, David Guennec, Aghilas Sini, Arnaud Delhay, Pierre-François Marteau |
INTERSPEECH | 5 |
| 2025 | Multi-level SSL Feature Gating for Audio Deepfake DetectionabstractRecent advancements in generative AI, particularly in speech synthesis, have enabled the generation of highly natural-sounding synthetic speech that closely mimics human voices. While these innovations hold promise for applications like assistive technologies, they also pose significant risks, including misuse for fraudulent activities, identity theft, and security threats. Current research on spoofing detection countermeasures remains limited by generalization to unseen deepfake attacks and languages. To address this, we propose a gating mechanism extracting relevant feature from the speech foundation XLS-R model as a front-end feature extractor. For downstream back-end classifier, we employ Multi-kernel gated Convolution (MultiConv) to capture both local and global speech artifacts. Additionally, we introduce Centered Kernel Alignment (CKA) as a similarity metric to enforce diversity in learned features across different MultiConv layers. By integrating CKA with our gating mechanism, we hypothesize that each component helps improving the learning of distinct synthetic speech patterns. Experimental results demonstrate that our approach achieves state-of-the-art performance on in-domain benchmarks while generalizing robustly to out-of-domain datasets, including multilingual speech samples. This underscores its potential as a versatile solution for detecting evolving speech deepfake threats. Hoan My Tran, Damien Lolive, Aghilas Sini, Arnaud Delhay, Pierre-François Marteau, David Guennec |
ACM Multimedia | 4 |
| 2024 | Spoofed Speech Detection with a Focus on Speaker EmbeddingabstractInternational audience Hoan My Tran, David Guennec, Aghilas Sini, Damien Lolive, Arnaud Delhay, Pierre-François Marteau |
INTERSPEECH | 6 |
| 2022 | Phone-Level Pronunciation Scoring for L1 Using Weighted-Dynamic Time WarpingabstractThis paper presents a novel approach for phone-level pronunciation scoring. The proposed method relies on the two usual stages of pronunciation scoring: an acoustic model transcribes the spoken utterance into a phoneme sequence and then, Weighted-Dynamic Time Warping (W-DTW) is used to compare the predicted phoneme sequence against the reference one. Our approach alters the comparison process by considering Phonetic PosteriorGrams (PPG) rather than only the most probable sequence of phonemes. This led us to propose a modified W-DTW algorithm that considers the probabilities of the predicted phonemes, as well as the use of articulatory features as a proxy of phonetic similarity. The results achieved are satisfactory considering the content of the adult speech database and are comparable to well-known state-of-the-art methods. Aghilas Sini, Antoine Perquin, Damien Lolive, Arnaud Delhay |
SLT | 4 |
| 2020 | Video Latent Code Interpolation for Anomalous Behavior DetectionabstractDetecting an anomalous human behavior can be a challenging task. In this paper, we present a novel objective function for autoencoders which include a temporal component. Our method is a fully end-to-end semi-supervised approach for video anomaly detection. The autoencoder is trained to reconstruct a sample from a partial input, by interpolating latent codes obtained from this partial input. We show this approach improves over using usual autoencoder objective functions for video anomaly detection and achieves results close to the state of the art on a broad range of datasets. Our code is publicly available on github. Valentin Durand de Gevigney, Pierre-François Marteau, Arnaud Delhay, Damien Lolive |
SMC | 3 |
| 2018 | EMO&LY (EMOtion and AnomaLY) : A new corpus for anomaly detection in an audiovisual stream with emotional context
Cédric Fayet, Arnaud Delhay, Damien Lolive, Pierre-François Marteau |
LREC | 2 |
| 2017 | Big Five vs. Prosodic Features as Cues to Detect Abnormality in SSPNET-Personality CorpusabstractThis paper presents an attempt to evaluate three different sets of features extracted from prosodic descriptors and Big Five traits for building an anomaly detector. The Big Five model enables to capture personality information. Big Five traits are extracted from a manual annotation while Prosodic features are extracted directly from the speech signal. Two different anomaly detection methods are evaluated: Gaussian Mixture Model (GMM) and One-Class SVM (OC-SVM), each one combined with a threshold classification to decide the ”normality” of a sample.
The different combinations of models and feature sets are evaluated on the SSPNET-Personality corpus which has already been
used in several experiments, including a previous work on separating two types of personality profiles in a supervised way.
In this work, we propose the above mentioned unsupervised or semi-supervised methods, and discuss their performance, to detect
particular audio-clips produced by a speaker with an abnormal personality. Results show that using automatically extracted
prosodic features competes with the Big Five traits. The overall detection performance achieved by the best model is
around 0.8 (F1-measure) Cédric Fayet, Arnaud Delhay, Damien Lolive, Pierre-François Marteau |
INTERSPEECH | 2 |
| 2015 | Large Linguistic Corpus Reduction with SCP AlgorithmsabstractLinguistic corpus design is a critical concern for building rich annotated corpora useful in different domains of applications. For example, speech technologies such as ASR (Automatic Speech Recognition) or TTS (Text-to-Speech) need a huge amount of speech data to train data-driven models or to produce synthetic speech. Collecting data is always related to costs (recording speech, verifying annotations, etc.), and as a rule of thumb, the more data you gather, the more costly your application will be. Within this context, we present in this article solutions to reduce the amount of linguistic text content while maintaining a sufficient level of linguistic richness required by a model or an application. This problem can be formalized as a Set Covering Problem (SCP) and we evaluate two algorithmic heuristics applied to design large text corpora in English and French for covering phonological information or POS labels. The first considered algorithm is a standard greedy solution with an agglomerative/spitting strategy and we propose a second algorithm based on Lagrangian relaxation. The latter approach provides a lower bound to the cost of each covering solution. This lower bound can be used as a metric to evaluate the quality of a reduced corpus whatever the algorithm applied. Experiments show that a suboptimal algorithm like a greedy algorithm achieves good results; the cost of its solutions is not so far from the lower bound (about 4.35% for 3-phoneme coverings). Usually, constraints in SCP are binary; we proposed here a generalization where the constraints on each covering feature can be multi-valued. Nelly Barbot, Olivier Boëffard, Jonathan Chevelu, Arnaud Delhay |
Comput. Linguistics | 4 |
| 2012 | Comparing performance of different set-covering strategies for linguistic content optimization in speech corpora
Nelly Barbot, Olivier Boëffard, Arnaud Delhay |
LREC | 3 |
| 2011 | Towards a Versatile Multi-Layered Description of Speech Corpora Using Algebraic RelationsabstractThis paper presents a software library, namely ROOTS for Rich Object Oriented Transcription System, that helps to describe spoken messages in a coherent manner linking sequences of items on numerous levels (linguistic, phonological, or acoustic). The proposed representation is incremental and can thus describe any or all parts of an utterance. In order to link different levels of description, algebraic relations are used. Instead of relying solely on fixed, pre-determined relations, algebraic composition operators are proposed that can create a missing relation on demand. In terms of software architecture, object classes are defined based on a well-grounded theoretical representation of speech (text, syntax, phonology and acoustics), without particular dependences on an annotation system (e.g. IPA is fully implemented). The API documentation for this software is available online [7]. Nelly Barbot, Vincent Barreaud, Olivier Boëffard, Laure Charonnat, Arnaud Delhay, Sébastien Le Maguer, Damien Lolive |
INTERSPEECH | 5 |
| 2008 | Comparing Set-Covering Strategies for Optimal Corpus Design
Jonathan Chevelu, Nelly Barbot, Olivier Boëffard, Arnaud Delhay |
LREC | 4 |
| 2008 | Analogical Dissimilarity: Definition, Algorithms and Two Experiments in Machine LearningabstractThis paper defines the notion of analogical dissimilarity between four objects, with a special focus on objects structured as sequences. Firstly, it studies the case where the four objects have a null analogical dissimilarity, i.e. are in analogical proportion. Secondly, when one of these objects is unknown, it gives algorithms to compute it. Thirdly, it tackles the problem of defining analogical dissimilarity, which is a measure of how far four objects are from being in analogical proportion. In particular, when objects are sequences, it gives a definition and an algorithm based on an optimal alignment of the four sequences. It gives also learning algorithms, i.e. methods to find the triple of objects in a learning sample which has the least analogical dissimilarity with a given object. Two practical experiments are described: the first is a classification problem on benchmarks of binary and nominal data, the second shows how the generation of sequences by solving analogical equations enables a handwritten character recognition system to rapidly be adapted to a new writer. Laurent Miclet, Sabri Bayoudh, Arnaud Delhay |
J. Artif. Intell. Res. | 3 |
| 2007 | Learning by Analogy: A Classification Rule for Binary and Nominal Data
Sabri Bayoudh, Laurent Miclet, Arnaud Delhay |
IJCAI | 3 |
| 1999 | Maximization of the Average Quality of Anytime Contract Algorithms over a Time Interval
Arnaud Delhay, Max Dauchet, Patrick Taillibert, Philippe Vanheeghe |
IJCAI | 1 |