David Doukhan

dblp:35/10650 · DBLP profile ↗
← Back
18ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0002-1645-7334ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 spINAch: A Diachronic Corpus of French Broadcast Speech Controlled for Speakers' Age and Gender
Simon Devauchelle, David Doukhan, Rémi Uro, Lucas Ondel Yang, Valentin Pelloin, Olympia Imbert-Brégégère, Véronique Lefort, Kévin Picard, Emeline Seignobos, Albert Rilliard
LREC2
2026 Data Selection Effects on Self-Supervised Learning of Audio Representations for French Audiovisual Broadcasts
Valentin Pelloin, Lina Bekkali, Réda Dehak, David Doukhan
LREC4
2024 InaGVAD : A Challenging French TV and Radio Corpus Annotated for Speech Activity Detection and Speaker Gender Segmentation
abstract
InaGVAD is an audio corpus collected from 10 French radio and 18 TV channels categorized into 4 groups: generalist radio, music radio, news TV, and generalist TV. It contains 277 1-minute-long annotated recordings aimed at representing the acoustic diversity of French audiovisual programs and was primarily designed to build systems able to monitor men’s and women’s speaking time in media. inaGVAD is provided with Voice Activity Detection (VAD) and Speaker Gender Segmentation (SGS) annotations extended with overlap, speaker traits (gender, age, voice quality), and 10 non-speech event categories. Annotation distributions are detailed for each channel category. This dataset is partitioned into a 1h development and a 3h37 test subset, allowing fair and reproducible system evaluation. A benchmark of 6 freely available VAD software is presented, showing diverse abilities based on channel and non-speech event categories. Two existing SGS systems are evaluated on the corpus and compared against a baseline X-vector transfer learning strategy, trained on the development subset. Results demonstrate that our proposal, trained on a single - but diverse - hour of data, achieved competitive SGS results. The entire inaGVAD package; including corpus, annotations, evaluation scripts, and baseline training code; is made freely accessible, fostering future advancement in the domain.
David Doukhan, Christine Maertens, William Le Personnic, Ludovic Speroni, Réda Dehak
LREC/COLING1
2024 Annotation of Transition-Relevance Places and Interruptions for the Description of Turn-Taking in Conversations in French Media Content
abstract
Few speech resources describe interruption phenomena, especially for TV and media content. The description of these phenomena may vary across authors: it thus leaves room for improved annotation protocols. We present an annotation of Transition-Relevance Places (TRP) and Floor-Taking event types on an existing French TV and Radio broadcast corpus to facilitate studies of interruptions and turn-taking. Each speaker change is annotated with the presence or absence of a TRP, and a classification of the next-speaker floor-taking as Smooth, Backchannel or different types of turn violations (cooperative or competitive, successful or attempted interruption). An inter-rater agreement analysis shows such annotations’ moderate to substantial reliability. The inter-annotator agreement for TRP annotation reaches κ=0.75, κ=0.56 for Backchannel and κ=0.5 for the Interruption/non-interruption distinction. More precise differences linked to cooperative or competitive behaviors lead to lower agreements. These results underline the importance of low-level features like TRP to derive a classification of turn changes that would be less subject to interpretation. The analysis of the presence of overlapping speech highlights the existence of interruptions without overlaps and smooth transitions with overlaps. These annotations are available at https://lium.univ-lemans.fr/corpus-allies/.
Rémi Uro, Marie Tahon, Jane Wottawa, David Doukhan, Albert Rilliard, Antoine Laurent
LREC/COLING4
2024 Gender Representation in TV and Radio: Automatic Information Extraction methods versus Manual Analyses
abstract
International audience
David Doukhan, Lena Dodson, Manon Conan, Valentin Pelloin, Aurélien Clamouse, Mélina Lepape, Géraldine Van Hille, Cécile Méadel, Marlène Coulomb-Gully
INTERSPEECH1
2024 Articulatory Configurations across Genders and Periods in French Radio and TV archives
abstract
International audience
Benjamin Elie, David Doukhan, Rémi Uro, Lucas Ondel Yang, Albert Rilliard, Simon Devauchelle
INTERSPEECH2
2024 Automatic Classification of News Subjects in Broadcast News: Application to a Gender Bias Representation Analysis
abstract
Automatic Classification of News Subjects in Broadcast News: Application to a Gender Bias Representation Analysis About Source code for the Interpseech 2024 paper about automatic classifcation of news subjects in broadcast news. The code contains evaluation scripts, models training and interfence source. Dataset The annotated dataset contains about 03h44min of broadcast news, with 605 dialogues for the Test set. The dataset can be downloaded at the following URL: https://www.ina.fr/recherche/dataset-project, under the name is24_news_topic. Evaluation scripts The scripts to eval your own models are included in the folder evaluation/. You will need an original copy of the dataset (see above.). You may use the following script: evaluation/eval \ --reference ORIGINAL_DATASET_FOLDER \ --prediction PRED_FOLDER/predictions.json \ --subset dev # dev or test The results will be saved inside PRED_FOLDER folder, with the name results.json. You need to format your predictions into the following format: { "DIALOGUE_ID_1": { "text": "input text for that dialogue (this field is not required in your output file, but it allow you to easily browse the file while reading the predicted output, and manually see if what your model predicted)", "class__ARTS/CULTURE/ENTERTAINMENT": false, "class__COMMERCIAL": false, "class__CRIME/LAW/JUSTICE": true, "class__DISASTER/ACCIDENT": false, "class__ECONOMY/BUSINESS/FINANCE": false, "class__EDUCATION": false, "class__ENVIRONMENTAL_ISSUE": false, "class__HEALTH": false, "class__LABOUR": false, "class__LIFESTYLE/LEISURE": false, "class__OTHER": false, "class__POLITICS": false, "class__RELIGION/BELIEF": false, "class__SCIENCE/TECHNOLOGY": false, "class__SOCIAL_ISSUE": false, "class__SPORT": false, "class__UNREST/CONFLICTS/WAR": true, "class__WEATHER": false }, "DIALOGUE_ID_2": { "..." } } Models The source code of training and inference of the models presented in the paper are included in the folder models/ (models/BERT and models/Mixtral). BERT finetuning The script to finetune all models presented in the paper is in models/BERT/train_all. It calls the train script. To generate the predictions of a specific model, use models/BERT/predict. You can then use evaluation/eval to get metric scores of your model. Mixtral The prompt used for the Mixtral model is in the file models/Mixtral/prompt.txt. Use the script models/Mixtral/inference_mixtral to generate the outputs of Mixtral (or an other model) base on this prompt and input dialogues. Then, call parse_safe_json to convert the output of the model into a valid JSON format, which you can then use with generate_prediction_dataset to create a valid training/validation/testing dataset. You then can use evaluation/eval to evaluate Mixtral response if the inference was on the dataset dialogues, or models/BERT/train to finetune BERT on these annotations. Installation You need python3, pip3. pip3 install -r requirements.txt Citation If you use this corpus or the source code of this repository, please cite the following article: @inproceedings{pelloin2024automatic, title = {Automatic Classification of News Subjects in Broadcast News: Application to a Gender Bias Representation Analysis}, author = {Valentin Pelloin and Lena Dodson and \'Emile Chapuis and Nicolas Hervé and David Doukhan}, booktitle = {Proc. InterSpeech 2024}, month = 9, year = 2024, address = "Kos Island, Greece", }
Valentin Pelloin, Lena Dodson, Emile Chapuis, Nicolas Hervé, David Doukhan
INTERSPEECH5
2024 Detecting the terminality of speech-turn boundary for spoken interactions in French TV and Radio content
abstract
International audience
Rémi Uro, Marie Tahon, David Doukhan, Antoine Laurent, Albert Rilliard
INTERSPEECH3
2023 Voice Passing : a Non-Binary Voice Gender Prediction System for evaluating Transgender voice transition
abstract
International audience
David Doukhan, Simon Devauchelle, Lucile Girard-Monneron, Mía Chávez Ruz, V. Chaddouk, Isabelle Wagner, Albert Rilliard
INTERSPEECH1
2022 A Semi-Automatic Approach to Create Large Gender- and Age-Balanced Speaker Corpora: Usefulness of Speaker Diarization & Identification
abstract
This paper presents a semi-automatic approach to create a diachronic corpus of voices balanced for speaker’s age, gender, and recording period, according to 32 categories (2 genders, 4 age ranges and 4 recording periods). Corpora were selected at French National Institute of Audiovisual (INA) to obtain at least 30 speakers per category (a total of 960 speakers; only 874 have be found yet). For each speaker, speech excerpts were extracted from audiovisual documents using an automatic pipeline consisting of speech detection, background music and overlapped speech removal and speaker diarization, used to present clean speaker segments to human annotators identifying target speakers. This pipeline proved highly effective, cutting down manual processing by a factor of ten. Evaluation of the quality of the automatic processing and of the final output is provided. It shows the automatic processing compare to up-to-date process, and that the output provides high quality speech for most of the selected excerpts. This method is thus recommendable for creating large corpora of known target speakers.
Rémi Uro, David Doukhan, Albert Rilliard, Laetitia Larcher, Anissa-Claire Adgharouamane, Marie Tahon, Antoine Laurent
LREC2
2021 Speaker Embeddings for Diarization of Broadcast Data In The Allies Challenge
abstract
Diarization consists in the segmentation of speech signals and the clustering of homogeneous speaker segments. State-of-the-art systems typically operate upon speaker embeddings, such as i-vectors or neural x-vectors, extracted from mel cepstral coefficients (MFCCs) or spectrograms. The recent SincNet architecture extracts x-vectors directly from raw speech signals. The work reported in this paper compares the performance of different embeddings extracted from MFCCs or the raw signal for speaker diarization and broadcast media treated with compression and sub-sampling, operations which typically degrade performance. Experiments are performed with the new ALLIES database that was designed to complement existing, publicly available French corpora of broadcast radio and TV shows. Results show that, in adverse conditions, with compression and sampling mismatch, SincNet x-vectors outperform i-vectors and x-vectors by relative DERs of 43% and 73% respectively. Additionally we found that SincNet x-vectors are not the absolute best embeddings but are more robust to data mismatch than others.
Anthony Larcher, Ambuj Mehrish, Marie Tahon, Sylvain Meignier, Jean Carrive, David Doukhan, Olivier Galibert, Nicholas W. D. Evans
ICASSP6
2018 An Open-Source Speaker Gender Detection Framework for Monitoring Gender Equality
abstract
This paper presents an approach based on acoustic analysis to describe gender equality in French audiovisual streams, through the estimation of male and female speaking time. Gender detection systems based on Gaussian Mixture Models, i-vectors and Convolutional Neural Networks (CNN) were trained using an internal database of 2,284 French speakers and evaluated using REPERE challenge corpus. The CNN system obtained the best performance with a frame-level gender detection F-measure of 96.52 and a hourly women speaking time percentage error bellow 0.6%. It was considered reliable enough to realize large-scale gender equality descriptions. The proposed gender detection system has been packaged as an open-source framework.
David Doukhan, Jean Carrive, Félicien Vallet, Anthony Larcher, Sylvain Meignier
ICASSP1
2018 Computer-assisted Speaker Diarization: How to Evaluate Human Corrections
Pierre-Alexandre Broux, David Doukhan, Simon Petit-Renaud, Sylvain Meignier, Jean Carrive
LREC2
2015 Analysing rhythm in ritual discourse in yucatec maya using automatic speech alignment
abstract
International audience
Valentina Vapnarsky, Claude Barras, Cédric Becquey, David Doukhan, Martine Adda-Decker, Lori Lamel
INTERSPEECH4
2012 Modelling pause duration as a function of contextual length
abstract
International audience
David Doukhan, Albert Rilliard, Sophie Rosset, Christophe d'Alessandro
INTERSPEECH1
2012 Designing French Tale Corpora for Entertaining Text To Speech Synthesis
David Doukhan, Sophie Rosset, Albert Rilliard, Christophe d'Alessandro, Martine Adda-Decker
LREC1
2011 Prosodic Analysis of a Corpus of Tales
abstract
International audience
David Doukhan, Albert Rilliard, Sophie Rosset, Martine Adda-Decker, Christophe d'Alessandro
INTERSPEECH1
2009 A High-Throughput Screening Approach to Discovering Good Forms of Biologically Inspired Visual Representation
abstract
While many models of biological object recognition share a common set of "broad-stroke" properties, the performance of any one model depends strongly on the choice of parameters in a particular instantiation of that model--e.g., the number of units per layer, the size of pooling kernels, exponents in normalization operations, etc. Since the number of such parameters (explicit or implicit) is typically large and the computational cost of evaluating one particular parameter set is high, the space of possible model instantiations goes largely unexplored. Thus, when a model fails to approach the abilities of biological visual systems, we are left uncertain whether this failure is because we are missing a fundamental idea or because the correct "parts" have not been tuned correctly, assembled at sufficient scale, or provided with enough training. Here, we present a high-throughput approach to the exploration of such parameter sets, leveraging recent advances in stream processing hardware (high-end NVIDIA graphic cards and the PlayStation 3's IBM Cell Processor). In analogy to high-throughput screening approaches in molecular biology and genetics, we explored thousands of potential network architectures and parameter instantiations, screening those that show promising object recognition performance for further analysis. We show that this approach can yield significant, reproducible gains in performance across an array of basic object recognition tasks, consistently outperforming a variety of state-of-the-art purpose-built vision systems from the literature. As the scale of available computational power continues to expand, we argue that this approach has the potential to greatly accelerate progress in both artificial vision and our understanding of the computational underpinning of biological vision.
Nicolas Pinto, David Doukhan, James J. DiCarlo, David D. Cox
PLoS Comput. Biol.2