Damien Lolive

dblp:77/2395 · DBLP profile ↗
← Back
34ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0002-1110-5444ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 2 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 2 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Automatic Prediction of Prominence and Boundary Strength from Text
abstract
International audience
Pauline Mas, Kévin Vythelingum, Jonathan Chevelu, Marion Ouédraogo, Damien Lolive, Olivier Rosec
LREC5
2025 Paraphrase Generation Evaluation Powered by an LLM: A Semantic Metric, Not a Lexical One
abstract
Evaluating automatic paraphrase production systems is a difficult task as it involves, among other things, assessing the semantic proximity between two sentences. Usual measures are based on lexical distances, or at least on semantic embedding alignments. The rise of Large Language Models (LLM) has provided tools to model relationships within a text thanks to the attention mechanism. In this article, we introduce ParaPLUIE, a new measure based on a log likelihood ratio from an LLM, to assess the quality of a potential paraphrase. This measure is compared with usual measures on two known by the NLP community datasets prior to this study. Three new small datasets have been built to allow metrics to be compared in different scenario and to avoid data contamination bias. According to evaluations, the proposed measure is better for sorting pairs of sentences by semantic proximity. In particular, it is much more independent to lexical distance and provides an interpretable classification threshold between paraphrases and non-paraphrases.
Quentin Lemesle, Jonathan Chevelu, Damien Lolive, Arnaud Delhay, Nelly Barbot
COLING4
2025 Audio Deepfake Source Tracing using Multi-Attribute Open-Set Identification and Verification
abstract
International audience
Pierre Falez, Tony Marteau, Damien Lolive, Arnaud Delhay
INTERSPEECH3
2025 Enabling the replicability of speech synthesis perceptual evaluations
abstract
How speech synthesis is evaluated is nowadays questioned. Not only have conventional listening tests as a whole been proven a poor match for modern synthesis, but more fundamentally, important information (e.g., the question asked to the listener) is frequently missing in the report of the outcome of the evaluation despite the impact on the interpretation of the test results. This can lead to uncertainty about the validity of these evaluations. To address this issue, we propose standardising the structure of any evaluation report. To facilitate this standardisation, our contribution is twofold: an open-source subjective evaluation platform; and a set of reporting guidelines. The platform is designed to enable the development of easily shareable evaluation recipes. The set of guidelines complements the platform to support researchers in reporting their evaluation choices and analysis in more detail while relying on the recipe to describe the actual evaluation process.
Sébastien Le Maguer, Gwénolé Lecorvé, Damien Lolive, Naomi Harte, Juraj Simko
INTERSPEECH3
2025 Leveraging SSL Speech Features and Mamba for Enhanced DeepFake Detection
abstract
International audience
Hoan My Tran, Damien Lolive, David Guennec, Aghilas Sini, Arnaud Delhay, Pierre-François Marteau
INTERSPEECH2
2025 Multi-level SSL Feature Gating for Audio Deepfake Detection
abstract
Recent advancements in generative AI, particularly in speech synthesis, have enabled the generation of highly natural-sounding synthetic speech that closely mimics human voices. While these innovations hold promise for applications like assistive technologies, they also pose significant risks, including misuse for fraudulent activities, identity theft, and security threats. Current research on spoofing detection countermeasures remains limited by generalization to unseen deepfake attacks and languages. To address this, we propose a gating mechanism extracting relevant feature from the speech foundation XLS-R model as a front-end feature extractor. For downstream back-end classifier, we employ Multi-kernel gated Convolution (MultiConv) to capture both local and global speech artifacts. Additionally, we introduce Centered Kernel Alignment (CKA) as a similarity metric to enforce diversity in learned features across different MultiConv layers. By integrating CKA with our gating mechanism, we hypothesize that each component helps improving the learning of distinct synthetic speech patterns. Experimental results demonstrate that our approach achieves state-of-the-art performance on in-domain benchmarks while generalizing robustly to out-of-domain datasets, including multilingual speech samples. This underscores its potential as a versatile solution for detecting evolving speech deepfake threats.
Hoan My Tran, Damien Lolive, Aghilas Sini, Arnaud Delhay, Pierre-François Marteau, David Guennec
ACM Multimedia2
2024 Spoofed Speech Detection with a Focus on Speaker Embedding
abstract
International audience
Hoan My Tran, David Guennec, Aghilas Sini, Damien Lolive, Arnaud Delhay, Pierre-François Marteau
INTERSPEECH5
2023 Local or Global: The Variation in the Encoding of Style Across Sentiment and Formality
Somayeh Jafaritazehjani, Gwénolé Lecorvé, Damien Lolive, John D. Kelleher
ICANN (10)3
2023 An extension of disentanglement metrics and its application to voice
abstract
International audience
Olivier Zhang, Olivier Le Blouch, Nicolas Gengembre, Damien Lolive
INTERSPEECH4
2022 Dispeech: A Synthetic Toy Dataset for Speech Disentangling
abstract
Recently, a growing interest in unsupervised learning of disentangled representations has been observed, with successful applications to both synthetic and real data. In speech processing, such methods have been able to disentangle speakers’ attributes from verbal content. To have a better understanding of disentanglement, synthetic data is necessary, as it provides a controllable framework to train models and evaluate disentanglement. Thus, we introduce diSpeech, a corpus of speech synthesized with the Klatt synthesizer. Its first version is constrained to vowels synthesized with 5 generative factors relying on pitch and formants. Experiments show the ability of variational autoencoders to disentangle these generative factors and assess the reliability of disentanglement metrics. In addition to provide a support to benchmark speech disentanglement methods, diSpeech also enables the objective evaluation of disentanglement on real speech, which is to our knowledge unprecedented. To illustrate this methodology, we apply it to TIMIT’s isolated vowels.
Olivier Zhang, Nicolas Gengembre, Olivier Le Blouch, Damien Lolive
ICASSP4
2022 A Low-Cost Motion Capture Corpus in French Sign Language for Interpreting Iconicity and Spatial Referencing Mechanisms
abstract
The automatic translation of sign language videos into transcribed texts is rarely approached in its whole, as it implies to finely model the grammatical mechanisms that govern these languages. The presented work is a first step towards the interpretation of French sign language (LSF) by specifically targeting iconicity and spatial referencing. This paper describes the LSF-SHELVES corpus as well as the original technology that was designed and implemented to collect it. Our goal is to use deep learning methods to circumvent the use of models in spatial referencing recognition. In order to obtain training material with sufficient variability, we designed a light-weight (and low-cost) capture protocol that enabled us to collect data from a large panel of LSF signers. This protocol involves the use of a portable device providing a 3D skeleton, and of a software developed specifically for this application to facilitate the post-processing of handshapes. The LSF-SHELVES includes simple and compound iconic and spatial dynamics, organized in 6 complexity levels, representing a total of 60 sequences signed by 15 LSF signers.
Clémence Mertz, Vincent Barreaud, Thibaut Le Naour, Damien Lolive, Sylvie Gibet
LREC4
2022 Investigating Inter- and Intra-speaker Voice Conversion using Audiobooks
abstract
Audiobook readers play with their voices to emphasize some text passages, highlight discourse changes or significant events, or in order to make listening easier and entertaining. A dialog is a central passage in audiobooks where the reader applies significant voice transformation, mainly prosodic modifications, to realize character properties and changes. However, these intra-speaker modifications are hard to reproduce with simple text-to-speech synthesis. The manner of vocalizing characters involved in a given story depends on the text style and differs from one speaker to another. In this work, this problem is investigated through the prism of voice conversion. We propose to explore modifying the narrator’s voice to fit the context of the story, such as the character who is speaking, using voice conversion. To this end, two complementary experiments are designed: the first one aims to assess the quality of our Phonetic PosteriorGrams (PPG)-based voice conversion system using parallel data. Subjective evaluations with naive raters are conducted to estimate the quality of the signal generated and the speaker similarity. The second experiment applies an intra-speaker voice conversion, considering narration passages and direct speech passages as two distinct speakers. Data are then nonparallel and the dissimilarity between character and narrator is subjectively measured.
Aghilas Sini, Damien Lolive, Nelly Barbot, Pierre Alain
LREC2
2022 Phone-Level Pronunciation Scoring for L1 Using Weighted-Dynamic Time Warping
abstract
This paper presents a novel approach for phone-level pronunciation scoring. The proposed method relies on the two usual stages of pronunciation scoring: an acoustic model transcribes the spoken utterance into a phoneme sequence and then, Weighted-Dynamic Time Warping (W-DTW) is used to compare the predicted phoneme sequence against the reference one. Our approach alters the comparison process by considering Phonetic PosteriorGrams (PPG) rather than only the most probable sequence of phonemes. This led us to propose a modified W-DTW algorithm that considers the probabilities of the predicted phonemes, as well as the use of articulatory features as a proxy of phonetic similarity. The results achieved are satisfactory considering the content of the adult speech database and are comparable to well-known state-of-the-art methods.
Aghilas Sini, Antoine Perquin, Damien Lolive, Arnaud Delhay
SLT3
2021 Neural-Driven Search-Based Paraphrase Generation
abstract
We study a search-based paraphrase generation scheme where candidate paraphrases are generated by iterated transformations from the original sentence and evaluated in terms of syntax quality, semantic distance, and lexical distance.The semantic distance is derived from BERT, and the lexical quality is based on GPT2 perplexity.To solve this multi-objective search problem, we propose two algorithms: Monte-Carlo Tree Search For Paraphrase Generation (MCPG) and Pareto Tree Search (PTS).We provide an extensive set of experiments on 5 datasets with a rigorous reproduction and validation for several state-of-the-art paraphrase generation algorithms.These experiments show that, although being non explicitly supervised, our algorithms perform well against these baselines.
Betty Fabre, Tanguy Urvoy, Jonathan Chevelu, Damien Lolive
EACL4
2021 Style as Sentiment Versus Style as Formality: The Same or Different?
Somayeh Jafaritazehjani, Gwénolé Lecorvé, Damien Lolive, John D. Kelleher
ICANN (5)3
2020 Style versus Content: A distinction without a (learnable) difference?
abstract
Textual style transfer involves modifying the style of a text while preserving its content. This assumes that it is possible to separate style from content. This paper investigates whether this separation is possible. We use sentiment transfer as our case study for style transfer analysis. Our experimental methodology frames style transfer as a multi-objective problem, balancing style shift with content preservation and fluency. Due to the lack of parallel data for style transfer we employ a variety of adversarial encoder-decoder networks in our experiments. Also, we use of a probing methodology to analyse how these models encode style-related features in their latent spaces. The results of our experiments which are further confirmed by a human evaluation reveal the inherent trade-off between the multiple style transfer objectives which indicates that style cannot be usefully separated from content within these style-transfer systems.
Somayeh Jafaritazehjani, Gwénolé Lecorvé, Damien Lolive, John D. Kelleher
COLING3
2020 Video Latent Code Interpolation for Anomalous Behavior Detection
abstract
Detecting an anomalous human behavior can be a challenging task. In this paper, we present a novel objective function for autoencoders which include a temporal component. Our method is a fully end-to-end semi-supervised approach for video anomaly detection. The autoencoder is trained to reconstruct a sample from a partial input, by interpolating latent codes obtained from this partial input. We show this approach improves over using usual autoencoder objective functions for video anomaly detection and achieves results close to the state of the art on a broad range of datasets. Our code is publicly available on github.
Valentin Durand de Gevigney, Pierre-François Marteau, Arnaud Delhay, Damien Lolive
SMC4
2020 Can We Generate Emotional Pronunciations for Expressive Speech Synthesis?
abstract
In the field of expressive speech synthesis, a lot of work has been conducted on suprasegmental prosodic features while few has been done on pronunciation variants. However, prosody is highly related to the sequence of phonemes to be expressed. This article raises two issues in the generation of emotional pronunciations for TTS systems. The first issue consists in designing an automatic pronunciation generation method from text, while the second issue addresses the very existence of emotional pronunciations through experiments conducted on emotional speech. To do so, an innovative pronunciation adaptation method which automatically adapts canonical phonemes first to those labeled in the corpus used to create a synthetic voice, then to those labeled in an expressive corpus, is presented. This method consists in training conditional random fields pronunciation models with prosodic, linguistic, phonological and articulatory features. The analysis of emotional pronunciations reveals strong dependencies between prosody and phoneme assimilation or elisions. According to perceptual tests, the double adaptation allows to synthesize expressive speech samples of good quality, but emotion-specific pronunciations are too subtle to be perceived by testers.
Marie Tahon, Gwénolé Lecorvé, Damien Lolive
IEEE Trans. Affect. Comput.3
2019 Corpus Design Using Convolutional Auto-Encoder Embeddings for Audio-Book Synthesis
abstract
International audience
Meysam Shamsi, Damien Lolive, Nelly Barbot, Jonathan Chevelu
INTERSPEECH2
2018 EMO&LY (EMOtion and AnomaLY) : A new corpus for anomaly detection in an audiovisual stream with emotional context
Cédric Fayet, Arnaud Delhay, Damien Lolive, Pierre-François Marteau
LREC3
2018 SynPaFlex-Corpus: An Expressive French Audiobooks Corpus dedicated to expressive speech synthesis
Aghilas Sini, Damien Lolive, Gaëlle Vidal, Marie Tahon, Elisabeth Delais-Roussarie
LREC2
2017 Big Five vs. Prosodic Features as Cues to Detect Abnormality in SSPNET-Personality Corpus
abstract
This paper presents an attempt to evaluate three different sets of features extracted from prosodic descriptors and Big Five traits for building an anomaly detector. The Big Five model enables to capture personality information. Big Five traits are extracted from a manual annotation while Prosodic features are extracted directly from the speech signal. Two different anomaly detection methods are evaluated: Gaussian Mixture Model (GMM) and One-Class SVM (OC-SVM), each one combined with a threshold classification to decide the ”normality” of a sample. The different combinations of models and feature sets are evaluated on the SSPNET-Personality corpus which has already been used in several experiments, including a previous work on separating two types of personality profiles in a supervised way. In this work, we propose the above mentioned unsupervised or semi-supervised methods, and discuss their performance, to detect particular audio-clips produced by a speaker with an abnormal personality. Results show that using automatically extracted prosodic features competes with the Big Five traits. The overall detection performance achieved by the best model is around 0.8 (F1-measure)
Cédric Fayet, Arnaud Delhay, Damien Lolive, Pierre-François Marteau
INTERSPEECH3
2016 On the Suitability of Vocalic Sandwiches in a Corpus-Based TTS Engine
abstract
Unit selection speech synthesis systems generally rely on target and concatenation costs for selecting the best unit sequence. The role of the concatenation cost is to insure that joining two voice segments will not cause any acoustic artefact to appear. For this task, acoustic distances (MFCC, F0) are typically used but in many cases, this is not enough to prevent concatenation artefacts. Among other strategies, the improvement of corpus covering by favouring units that naturally support well the joining process (vocalic sandwiches) seems to be effective on TTS. In this paper, we investigate if vocalic sandwiches can be used directly in the unit selection engine when the corpus was not created using that principle. First, the sandwich approach is directly transposed in the unit selection engine with a penalty that greatly favours concatenation on sandwich boundaries. Second, a derived fuzzy version is proposed to relax the penalty based on the concatenation cost, with respect to the cost distribution. We show that the sandwich approach, very efficient at the corpus creation step, seems to be inefficient when directly transposed in the unit selection engine. However, we observe that the fuzzy approach enhances synthesis quality, especially on sentences with high concatenation costs.
David Guennec, Damien Lolive
INTERSPEECH2
2016 Improving TTS with Corpus-Specific Pronunciation Adaptation
abstract
Text-to-speech (TTS) systems are built on speech corpora which are labeled with carefully checked and segmented phonemes. However, phoneme sequences generated by automatic grapheme-to-phoneme converters during synthesis are usually inconsistent with those from the corpus, thus leading to poor quality synthetic speech signals. To solve this problem , the present work aims at adapting automatically generated pronunciations to the corpus. The main idea is to train corpus-specific phoneme-to-phoneme conditional random fields with a large set of linguistic, phonological, articulatory and acoustic-prosodic features. Features are first selected in cross-validation condition, then combined to produce the final best feature set. Pronunciation models are evaluated in terms of phoneme error rate and through perceptual tests. Experiments carried out on a French speech corpus show an improvement in the quality of speech synthesis when pronunciation models are included in the phonetization process. Appart from improving TTS quality, the presented pronunciation adaptation method also brings interesting perspectives in terms of expressive speech synthesis.
Marie Tahon, Raheel Qader, Gwénolé Lecorvé, Damien Lolive
INTERSPEECH4
2015 Adaptive statistical utterance phonetization for French
abstract
Traditional utterance phonetization methods concatenate pronunciations of uncontextualized constituent words. This approach is too weak for some languages, like French, where transitions between words imply pronunciation modifications. Moreover, it makes it difficult to consider global pronunciation strategies, for instance to model a specific speaker or a specific accent. To overcome these problems, this paper presents a new original phonetization approach for French to generate pronunciation variants of utterances. This approach offers a statistical and highly adaptive framework by relying on conditional random fields and weighted finite state transducers. The approach is evaluated on a corpus of isolated words and a corpus of spoken utterances.
Gwénolé Lecorvé, Damien Lolive
ICASSP2
2015 How to compare TTS systems: a new subjective evaluation methodology focused on differences
abstract
International audience
Jonathan Chevelu, Damien Lolive, Sébastien Le Maguer, David Guennec
INTERSPEECH2
2014 Towards the adaptation of prosodic models for expressive text-to-speech synthesis
Mathieu Avanzi, George Christodoulides, Damien Lolive, Elisabeth Delais-Roussarie, Nelly Barbot
INTERSPEECH3
2014 Adapting prosodic chunking algorithm and synthesis system to specific style: the case of dictation
abstract
International audience
Elisabeth Delais-Roussarie, Damien Lolive, Hiyon Yoo, Nelly Barbot, Olivier Rosec
INTERSPEECH2
2014 ROOTS: a toolkit for easy, fast and consistent processing of large sequential annotated data collections
Jonathan Chevelu, Gwénolé Lecorvé, Damien Lolive
LREC3
2012 Towards Fully Automatic Annotation of Audio Books for TTS
Olivier Boëffard, Laure Charonnat, Sébastien Le Maguer, Damien Lolive
LREC4
2011 Towards a Versatile Multi-Layered Description of Speech Corpora Using Algebraic Relations
abstract
This paper presents a software library, namely ROOTS for Rich Object Oriented Transcription System, that helps to describe spoken messages in a coherent manner linking sequences of items on numerous levels (linguistic, phonological, or acoustic). The proposed representation is incremental and can thus describe any or all parts of an utterance. In order to link different levels of description, algebraic relations are used. Instead of relying solely on fixed, pre-determined relations, algebraic composition operators are proposed that can create a missing relation on demand. In terms of software architecture, object classes are defined based on a well-grounded theoretical representation of speech (text, syntax, phonology and acoustics), without particular dependences on an annotation system (e.g. IPA is fully implemented). The API documentation for this software is available online [7].
Nelly Barbot, Vincent Barreaud, Olivier Boëffard, Laure Charonnat, Arnaud Delhay, Sébastien Le Maguer, Damien Lolive
INTERSPEECH7
2009 An evaluation methodology for prosody transformation systems based on chirp signals
abstract
Evaluation of prosody transformation systems is an important issue. First, the existing evaluation methodologies focus on parallel evaluation of systems and are not applicable to compare parallel and non-parallel systems. Secondly, these methodologies do not guarantee the independence from other features such as the segmental component. In particular, its influence cannot be neglected during evaluation and introduces a bias in the listening test. To answer these problems, we propose an evaluation methodology that depends only on the melody of the voice and that is applicable in a non-parallel context. Given a melodic contour, we propose to build an audio whistle from a chirp signal model. Experimental results show the efficiency of the proposed method concerning the discrimination of voices using only their melody information. An example of transformation function is also given and the results confirm the applicability of this methodology. Index Terms: subjective evaluation, prosody transformation, chirp signal, non-parallel corpora
Damien Lolive, Nelly Barbot, Olivier Boëffard
INTERSPEECH1
2007 Unsupervised HMM classification of F0 curves
abstract
This article describes a new unsupervised methodology to learn F0 classes using HMM models on a syllable basis. A F0 class is represented by a HMM with three emitting states. The clustering algorithm relies on an iterative gaussian splitting and EM retraining process. First, a single class is learnt on a training corpus (8000 syllables) and it is then divided by perturbing gaussian means of successive levels. At each step, the mean RMS error is evaluated on a validation corpus (3000 syllables). The algorithm stops automatically when the error becomes stable or increases. The syllabic structure of a sentence is the reference level we have taken for F0 modelling even if the methodology can be applied to other structures. Clustering quality is evaluated in terms of cross-validation using a mean of RMS errors between F0 contours on a test corpus and the estimated HMM trajectories. The results show a pretty good quality of the classes (mean RMS error around 4Hz). Index Terms: prosody, fundamental frequency, unsupervised classification, Hidden Markov Model
Damien Lolive, Nelly Barbot, Olivier Boëffard
INTERSPEECH1
2005 F0 stylisation with a free-knot b-spline model and simulated-annealing optimization
abstract
International audience
Nelly Barbot, Olivier Boëffard, Damien Lolive
INTERSPEECH3