Benjamin O'Brien

dblp:230/6190 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
10since 2021 · last 2024
0000-0002-1255-8410ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Evaluating the effects of task design on unfamiliar Francophone listener and automatic speaker identification performance
Benjamin O'Brien, Christine Meunier, Natalia A. Tomashenko, Alain Ghio, Jean-François Bonastre
Multim. Tools Appl.1
2024 Evaluating the effects of continuous pitch and speech tempo modifications on perceptual speaker verification performance by familiar and unfamiliar listeners
Benjamin O'Brien, Christine Meunier, Alain Ghio
Speech Commun.1
2023 Enhancing Expressivity Transfer in Textless Speech-to-Speech Translation
abstract
Textless speech-to-speech translation systems are rapidly advancing, thanks to the integration of self-supervised learning techniques. However, existing state-of-the-art systems fall short when it comes to capturing and transferring expressivity accurately across different languages. Expressivity plays a vital role in conveying emotions, nuances, and cultural subtleties, thereby enhancing communication across diverse languages. To address this issue this study presents a novel method that operates at the discrete speech unit level and leverages multilingual emotion embeddings to capture language-agnostic information. Specifically, we demonstrate how these embeddings can be used to effectively predict the pitch and duration of speech units in the target language. Through objective and subjective experiments conducted on a French-to-English translation task, our findings highlight the superior expressivity transfer achieved by our approach compared to current state-of-the-art systems.
Jarod Duret, Benjamin O'Brien, Yannick Estève, Titouan Parcollet
ASRU2
2023 Differences between Mimicking and Non-Mimicking laughter in Child-Caregiver Conversation: A Distributional and Acoustic Analysis
Chiara Mazzocconi, Benjamin O'Brien, Kevin El Haddad, Kübra Bodur, Abdellah Fourtassi
CogSci2
2023 Describing the phonetics in the underlying speech attributes for deep and interpretable speaker recognition
abstract
International audience
Imen Ben Amor, Jean-François Bonastre, Benjamin O'Brien, Pierre-Michel Bousquet
INTERSPEECH3
2023 Differentiating acoustic and physiological features in speech for hypoxia detection
abstract
International audience
Benjamin O'Brien, Adrien Gresse, Jean-Baptise Billaud, Guilhem Belda, Jean-François Bonastre
INTERSPEECH1
2022 Evaluating the effects of modified speech on perceptual speaker identification performance
abstract
International audience
Benjamin O'Brien, Christine Meunier, Alain Ghio
INTERSPEECH1
2022 The VoicePrivacy 2020 Challenge: Results and findings
Natalia A. Tomashenko, Xin Wang 0037, Emmanuel Vincent 0001, Jose Patino 0001, Brij Mohan Lal Srivastava, Paul-Gauthier Noé, Andreas Nautsch, Nicholas W. D. Evans, Junichi Yamagishi, Benjamin O'Brien, Anaïs Chanclu, Jean-François Bonastre, Massimiliano Todisco, Mohamed Maouche
Comput. Speech Lang.10
2021 Presentation Matters: Evaluating Speaker Identification Tasks
Benjamin O'Brien, Christine Meunier, Alain Ghio
Interspeech1
2021 Anonymous Speaker Clusters: Making Distinctions Between Anonymised Speech Recordings with Clustering Interface
abstract
Our study examined the performance of evaluators tasked to group natural and anonymised speech recordings into clusters based on their perceived similarities. Speech stimuli were selected from the VCTK corpus; two systems developed for the VoicePrivacy 2020 Challenge were used for anonymisation. The Baseline-1 (B1) system was developed by using x-vectors and neural waveform models, while the Baseline-2 (B2) system relied on digital-signal-processing techniques. 74 evaluators completed three trials composed of 16 recordings with either natural or anonymised speech generated from a single system. F-measure and cluster purity metrics were used to assess evaluator accuracy. Probabilistic linear discriminant analysis (PLDA) scores from an automatic speaker verification system were generated to quantify similarity between recordings and used to correlate subjective results. Our findings showed that non-native English speaking evaluators significantly lowered their F-measure means when presented anonymised recordings. We observed no significance for cluster purity. Pearson correlation procedures revealed that PLDA scores generated from natural and B2-anonymised speech recordings correlated positively to F-measure and cluster purity metrics. These findings show evaluators were able to use the interface to cluster natural and anonymised speech recordings and suggest anonymisation systems modelled like B1 are more effective at suppressing identifiable speech characteristics.
Benjamin O'Brien, Natalia A. Tomashenko, Anaïs Chanclu, Jean-François Bonastre
Interspeech1