Sébastien Le Maguer

dblp:47/10649 · DBLP profile ↗
← Back
23ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0003-0407-8318ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 6 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 The use of variable length stimuli for assessing segmental distortion in TTS evaluation
Ayushi Pandey, Jens Edlund, Sébastien Le Maguer, Naomi Harte
Comput. Speech Lang.3
2026 ASVspoof 5: Design, collection and validation of resources for spoofing, deepfake, and adversarial attack detection using crowdsourced speech
abstract
ASVspoof 5 is the fifth edition in a series of challenges which promote the study of speech spoofing and deepfake attacks as well as the design of detection solutions. We introduce the ASVspoof 5 database which is generated in a crowdsourced fashion from data collected in diverse acoustic conditions (cf. studio-quality data for earlier ASVspoof databases) and from ∼ 2,000 speakers (cf. ∼ 100 earlier). The database contains attacks generated with 32 different algorithms, also crowdsourced, and optimised to varying degrees using new surrogate detection models. Among them are attacks generated with a mix of legacy and contemporary text-to-speech synthesis and voice conversion models, in addition to adversarial attacks which are incorporated for the first time. ASVspoof 5 protocols comprise seven speaker-disjoint partitions. They include two distinct partitions for the training of different sets of attack models, two more for the development and evaluation of surrogate detection models, and then three additional partitions which comprise the ASVspoof 5 training, development and evaluation sets. An auxiliary set of data collected from an additional 30k speakers can also be used to train speaker encoders for the implementation of attack algorithms. Also described herein is an experimental validation of the new ASVspoof 5 database using a set of automatic speaker verification and spoof/deepfake baseline detectors. With the exception of protocols and tools for the generation of spoofed/deepfake speech, the resources described in this paper, already used by participants of the ASVspoof 5 challenge in 2024, are now all freely available to the community.
Xin Wang 0037, Héctor Delgado, Hemlata Tak, Jee-Weon Jung, Hye-Jin Shim, Massimiliano Todisco, Ivan Kukanov, Xuechen Liu 0001, Md. Sahidullah, Tomi Kinnunen, Nicholas W. D. Evans, Kong-Aik Lee, Junichi Yamagishi, Myeonghun Jeong, Yongyi Zang, Soumi Maiti, Florian Lux, Nicolas Müller, Wangyou Zhang, Chengzhe Sun 0001, Shuwei Hou, Siwei Lyu, Sébastien Le Maguer, Hanjie Guo, Vishwanath Pratap Singh
Comput. Speech Lang.25
2025 Enabling the replicability of speech synthesis perceptual evaluations
abstract
How speech synthesis is evaluated is nowadays questioned. Not only have conventional listening tests as a whole been proven a poor match for modern synthesis, but more fundamentally, important information (e.g., the question asked to the listener) is frequently missing in the report of the outcome of the evaluation despite the impact on the interpretation of the test results. This can lead to uncertainty about the validity of these evaluations. To address this issue, we propose standardising the structure of any evaluation report. To facilitate this standardisation, our contribution is twofold: an open-source subjective evaluation platform; and a set of reporting guidelines. The platform is designed to enable the development of easily shareable evaluation recipes. The set of guidelines complements the platform to support researchers in reporting their evaluation choices and analysis in more detail while relying on the recipe to describe the actual evaluation process.
Sébastien Le Maguer, Gwénolé Lecorvé, Damien Lolive, Naomi Harte, Juraj Simko
INTERSPEECH1
2025 The Effect of Voice and Repair Strategy on Trust Formation and Repair in Human-Robot Interaction
abstract
Trust is essential for social interactions, including those between humans and social artificial agents, such as robots. Several factors and combinations thereof can contribute to the formation of trust and, importantly in the case of machines that work with a certain margin of error, to its maintenance and repair after it has been breached. In this article, we present the results of a study aimed at investigating the role of robot voice and chosen repair strategy on trust formation and repair in a collaborative task. People helped a robot navigate through a maze, and the robot made mistakes at pre-defined points during the navigation. Via in-game behaviour and follow-up questionnaires, we could measure people’s trust towards the robot. We found that people trusted the robot speaking with a state-of-the-art synthetic voice more than with the default robot voice in the game, even though they indicated the opposite in the questionnaires. Additionally, we found that three repair strategies that people use in human-human interaction (justification of the mistake, promise to be better and denial of the mistake) work also in human-robot interaction.
Marta Romeo, Ilaria Torre 0002, Sébastien Le Maguer, Alexander Sleat, Angelo Cangelosi, Iolanda Leite
ACM Trans. Hum. Robot Interact.3
2024 Assessing the impact of contextual framing on subjective TTS quality
abstract
Edlund J, Tånnander C, LeMaguer S, Wagner P. Assessing the impact of contextual framing on subjective TTS quality. In: Proceedings of INTERSPEECH 2024. 2024: 1205--1209.
Jens Edlund, Christina Tånnander, Sébastien Le Maguer, Petra Wagner
INTERSPEECH3
2024 The limits of the Mean Opinion Score for speech synthesis evaluation
Sébastien Le Maguer, Simon King 0001, Naomi Harte
Comput. Speech Lang.1
2023 Sp1NY: A Quick and Flexible Speech Visualisation Tool in Python
Sébastien Le Maguer, Mark Anderson 0006, Naomi Harte
INTERSPEECH1
2023 Listener sensitivity to deviating obstruents in WaveNet
Ayushi Pandey, Jens Edlund, Sébastien Le Maguer, Naomi Harte
INTERSPEECH3
2023 Putting Robots in Context: Challenging the Influence of Voice and Empathic Behaviour on Trust
abstract
Trust is essential for social interactions, including those between humans and social artificial agents, such as robots. Several robot-related factors can contribute to the formation of trust. However, previous work has often treated trust as an absolute concept, whereas it is highly context-dependent, and it is possible that some robot-related features will influence trust in some contexts, but not in others. In this paper, we present the results of two video-based online studies aimed at investigating the role of robot voice and empathic behaviour on trust formation in a general context as well as in a task-specific context. We found that voice influences trust in the specific context, with no effect of voice or empathic behaviour in the general context. Thus, context mediated whether robot-related features play a role in people’s trust formation towards robots.
Marta Romeo, Ilaria Torre 0002, Sébastien Le Maguer, Angelo Cangelosi, Iolanda Leite
RO-MAN3
2022 Robo-Identity: Exploring Artificial Identity and Emotion via Speech Interactions
abstract
Following the success of the first edition of Robo-Identity, the second edition will provide an opportunity to expand the discussion about artificial identity. This year, we are focusing on emotions that are expressed through speech and voice. Synthetic voices of robots can resemble and are becoming indistinguishable from expressive human voices. This can be an opportunity and a constraint in expressing emotional speech that can (falsely) convey a human-like identity that can mislead people, leading to ethical issues. How should we envision an agent's artificial identity? In what ways should we have robots that maintain a machine-like stance, e.g., through robotic speech, and should emotional expressions that are increasingly human-like be seen as design opportunities? These are not mutually exclusive concerns. As this discussion needs to be conducted in a multidisciplinary manner, we welcome perspectives on challenges and opportunities from variety of fields. For this year's edition, the special theme will be “speech, emotion and artificial identity”.
Guy Laban, Sébastien Le Maguer, Minha Lee, Dimosthenis Kontogiorgos, Samantha Reig, Ilaria Torre 0002, Ravi Tejwani, Matthew J. Dennis, André Pereira 0001
HRI2
2022 Back to the Future: Extending the Blizzard Challenge 2013
Sébastien Le Maguer, Simon King 0001, Naomi Harte
INTERSPEECH1
2022 Production characteristics of obstruents in WaveNet and older TTS systems
Ayushi Pandey, Sébastien Le Maguer, Julie Carson-Berndsen, Naomi Harte
INTERSPEECH2
2020 Can Auditory Nerve Models Tell us What's Different About WaveNet Vocoded Speech?
Sébastien Le Maguer, Naomi Harte
INTERSPEECH1
2020 Investigation of Auditory Nerve Model Based Analysis for Vocoded Speech Synthesis
abstract
In recent decades, the quality of speech synthesized by computers has increased drastically. However, evaluating such systems remains a challenge as the relevant methodologies haven't evolved for more than a decade. Subjective evaluation provides a global overview of the quality, but lacks any detailed feedback. Furthermore, research in objective evaluation hasn't yet delivered any detailed analysis methodologies. Inspired by the speech intelligibility and speech quality fields, we investigate how we can use an Auditory Nerve (AN) model to improve objective evaluation of speech synthesis systems. To do so, we compare different configurations of Hidden Markov Model (HMM) and deep neural network (DNN) synthesis using two different metrics derived from spectrograms, mean-rate neurograms and fine-timing neurograms. The metrics are the Root Mean Square Error (RMSE) and the Neurogram Similarity Index Measure (NSIM). As using an AN model introduces a perceptual angle in the analysis, we also compare the different configurations using two established perceptual-based quality models: Perceptual Evaluation of Speech Quality (PESQ) and Virtual Speech Quality Objective Listener (ViSQOL). The results show ViSQOL and PESQ are not suitable to a refined analysis of speech synthesis. The results also show that comparing mean-rate neurograms using the NSIM metric is an effective alternative to the comparison of spectrograms using the RMSE.
Sébastien Le Maguer, Naomi Harte
QoMEX1
2020 Should robots have accents?
abstract
Accents are vocal features that immediately tell a listener whether a speaker comes from their same place, i.e. whether they share a social group. This in-groupness is important, as people tend to prefer interacting with others who belong to their same groups. Accents also evoke attitudinal responses based on their supposed prestigious status. These accent-based perceptions might affect interactions between humans and robots. Yet, very few studies so far have investigated the effect of accented robot speakers on users' perceptions and behaviour, and none have collected users' explicit preferences on robot accents. In this paper we present results from a survey of over 500 British speakers, who indicated what accent they would like a robot to have. The biggest proportion of participants wanted a robot to have a Standard Southern British English (SSBE) accent, followed by an Irish accent. Crucially, very few people wanted a robot with their same accent, or with a machine-like voice. These explicit preferences might not turn out to predict more successful interactions, also because of the unrealistic expectations that such human-like vocal features might generate in a user. Nonetheless, it seems that people have an idea of how their artificial companions should sound like, and this preference should be considered when designing them.
Ilaria Torre 0002, Sébastien Le Maguer
RO-MAN2
2020 ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech
Xin Wang 0037, Junichi Yamagishi, Massimiliano Todisco, Héctor Delgado, Andreas Nautsch, Nicholas W. D. Evans, Md. Sahidullah, Ville Vestman, Tomi Kinnunen, Kong-Aik Lee, Lauri Juvela, Paavo Alku, Yu-Huai Peng, Hsin-Te Hwang, Yu Tsao 0001, Hsin-Min Wang, Sébastien Le Maguer, Zhen-Hua Ling
Comput. Speech Lang.17
2018 Creating New Language and Voice Components for the Updated MaryTTS Text-to-Speech Synthesis Platform
Ingmar Steiner, Sébastien Le Maguer
LREC2
2017 Shadowing Synthesized Speech - Segmental Analysis of Phonetic Convergence
Iona Gessinger, Eran Raveh, Sébastien Le Maguer, Bernd Möbius, Ingmar Steiner
INTERSPEECH3
2017 An HMM/DNN Comparison for Synchronized Text-to-Speech and Tongue Motion Synthesis
Sébastien Le Maguer, Ingmar Steiner, Alexander Hewer
INTERSPEECH1
2017 Synthesis of Tongue Motion and Acoustics From Text Using a Multimodal Articulatory Database
abstract
We present an end-to-end text-to-speech (TTS) synthesis system that generates audio and synchronized tongue motion directly from text. This is achieved by adapting a three-dimensional model of the tongue surface to an articulatory dataset and training a statistical parametric speech synthesis system directly on the tongue model parameters. We evaluate the model at every step by comparing the spatial coordinates of predicted articulatory movements against the reference data. The results indicate a global mean Euclidean distance of less than 2.8 mm, and our approach can be adapted to add an articulatory modality to conventional TTS applications without the need for extra data.
Ingmar Steiner, Sébastien Le Maguer, Alexander Hewer
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 How to compare TTS systems: a new subjective evaluation methodology focused on differences
abstract
International audience
Jonathan Chevelu, Damien Lolive, Sébastien Le Maguer, David Guennec
INTERSPEECH3
2012 Towards Fully Automatic Annotation of Audio Books for TTS
Olivier Boëffard, Laure Charonnat, Sébastien Le Maguer, Damien Lolive
LREC3
2011 Towards a Versatile Multi-Layered Description of Speech Corpora Using Algebraic Relations
abstract
This paper presents a software library, namely ROOTS for Rich Object Oriented Transcription System, that helps to describe spoken messages in a coherent manner linking sequences of items on numerous levels (linguistic, phonological, or acoustic). The proposed representation is incremental and can thus describe any or all parts of an utterance. In order to link different levels of description, algebraic relations are used. Instead of relying solely on fixed, pre-determined relations, algebraic composition operators are proposed that can create a missing relation on demand. In terms of software architecture, object classes are defined based on a well-grounded theoretical representation of speech (text, syntax, phonology and acoustics), without particular dependences on an annotation system (e.g. IPA is fully implemented). The API documentation for this software is available online [7].
Nelly Barbot, Vincent Barreaud, Olivier Boëffard, Laure Charonnat, Arnaud Delhay, Sébastien Le Maguer, Damien Lolive
INTERSPEECH6