VLDB 2026 Research / reviewers in the wild / expert
Ioana Vasilescu
dblp:24/2297
· DBLP profile ↗
39ranked-venue papers
9as first author
13since 2021 · last 2026
0000-0003-3038-4503ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 9 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 7 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Added Value of Metadata and Annotations: Evidence from Two Large-Scale, Naturalistic Corpus Studies
Anisia Popescu, Johanna Cronenberg, Ioana Vasilescu, Ioana Chitoran, Lori Lamel, Martine Adda-Decker |
LREC | 3 |
| 2025 | Tracking /r/ Deletion: Forced Alignment of Pronunciation Variants and Sociophonetic Insights into Post-Obstruent Final /r/ in FrenchabstractInternational audience Anisia Popescu, Lori Lamel, Marc Evrard, Ioana Vasilescu |
INTERSPEECH | 4 |
| 2024 | Linguistic Nudges and Verbal Interaction with Robots, Smart-Speakers, and HumansabstractThis paper describes a data collection methodology and emotion annotation of dyadic interactions between a human, a Pepper robot, a Google Home smart-speaker, or another human. The collected 16 hours of audio recordings were used to analyze the propensity to change someone’s opinions about ecological behavior regarding the type of conversational agent, the kind of nudges, and the speaker’s emotional state. We describe the statistics of data collection and annotation. We also report the first results, which showed that humans change their opinions on more questions with a human than with a device, even against mainstream ideas. We observe a correlation between a certain emotional state and the interlocutor and a human’s propensity to be influenced. We also reported the results of the studies that investigated the effect of human likeness on speech using our data. Natalia Kalashnikova, Ioana Vasilescu, Laurence Devillers |
LREC/COLING | 2 |
| 2024 | Using Speech Technology to Test Theories of Phonetic and Phonological TypologyabstractThe present paper uses speech technology derived tools and methodologies to test theories about phonetic typology. We specifically look at how the two-way laryngeal contrast (voiced /b, d, g, v, z/ vs. voiceless /p, t, k, f, s/ obstruents) is implemented in European Portuguese, a language that has been suggested to exhibit a different voicing system than its sister Romance languages, more similar to the one found for Germanic languages. A large European Portuguese corpus was force aligned using (1) different combinations of parallel Portuguese (original), Italian (Romance language) and German (Germanic language) acoustic phone models and letting an ASR system choose the best fitting one, and (2) pronunciation variants (/b, d, g, v, z/ produced as either [b, d, g, v, z] or [p, t, k, f, s]) for obstruent consonants. Results support previous accounts in the literature that European Portuguese is diverging from the traditional voicing system known for Romance language, towards a hybrid system where stops and fricatives are specified for different voicing features. Anisia Popescu, Lori Lamel, Ioana Vasilescu |
LREC/COLING | 3 |
| 2024 | Crosslinguistic Comparison of Acoustic Variation in the Vowel Sequences /ia/ and /io/ in Four Romance LanguagesabstractInternational audience Johanna Cronenberg, Ioana Chitoran, Lori Lamel, Ioana Vasilescu |
INTERSPEECH | 4 |
| 2024 | Automatic Speech Recognition with parallel L1 and L2 acoustic phone models to evaluate /l/ allophony in L2 English speech production
Anisia Popescu, Lori Lamel, Ioana Vasilescu, Laurence Devillers |
INTERSPEECH | 3 |
| 2022 | Eye Got It: A System for Automatic Calculation of the Eye-Voice Span
Mohamed El Baha, Olivier Augereau, Sofiya Kobylyanskaya, Ioana Vasilescu, Laurence Devillers |
DAS | 4 |
| 2022 | When Phonetics Meets Morphology: Intervocalic Voicing Within and Across Words in Romance LanguagesabstractInternational audience Mathilde Hutin, Martine Adda-Decker, Lori Lamel, Ioana Vasilescu |
INTERSPEECH | 4 |
| 2022 | Voicing neutralization in Romanian fricatives across different speech stylesabstractInternational audience Laura Spinu, Ioana Vasilescu, Lori Lamel, Jason Lilley |
INTERSPEECH | 2 |
| 2022 | Corpus Design for Studying Linguistic Nudges in Human-Computer Spoken InteractionsabstractIn this paper, we present the methodology of corpus design that will be used to study the comparison of influence between linguistic nudges with positive or negative influences and three conversational agents: robot, smart speaker, and human. We recruited forty-nine participants to form six groups. The conversational agents first asked the participants about their willingness to adopt five ecological habits and invest time and money in ecological problems. The participants were then asked the same questions but preceded by one linguistic nudge with positive or negative influence. The comparison of standard deviation and mean metrics of differences between these two notes (before the nudge and after) showed that participants were mainly affected by nudges with positive influence, even though several nudges with negative influence decreased the average note. In addition, participants from all groups were willing to spend more money than time on ecological problems. In general, our experiment’s early results suggest that a machine agent can influence participants to the same degree as a human agent. A better understanding of the power of influence of different conversational machines and the potential of influence of nudges of different polarities will lead to the development of ethical norms of human-computer interactions. Natalia Kalashnikova, Serge Pajak, Fabrice Le Guel, Ioana Vasilescu, Gemma Serrano, Laurence Devillers |
LREC | 4 |
| 2022 | Extracting Linguistic Knowledge from Speech: A Study of Stop Realization in 5 Romance LanguagesabstractThis paper builds upon recent work in leveraging the corpora and tools originally used to develop speech technologies for corpus-based linguistic studies. We address the non-canonical realization of consonants in connected speech and we focus on voicing alternation phenomena of stops in 5 standard varieties of Romance languages (French, Italian, Spanish, Portuguese, Romanian). For these languages, both large scale corpora and speech recognition systems were available for the study. We use forced alignment with pronunciation variants and machine learning techniques to examine to what extent such frequent phenomena characterize languages and what are the most triggering factors. The results confirm that voicing alternations occur in all Romance languages. Automatic classification underlines that surrounding contexts and segment duration are recurring contributing factors for modeling voicing alternation. The results of this study also demonstrate the new role that machine learning techniques such as classification algorithms can play in helping to extract linguistic knowledge from speech and to suggest interesting research directions. Yaru Wu, Mathilde Hutin, Ioana Vasilescu, Lori Lamel, Martine Adda-Decker |
LREC | 3 |
| 2022 | Using a Knowledge Base to Automatically Annotate Speech Corpora and to Identify Sociolinguistic VariationabstractSpeech characteristics vary from speaker to speaker. While some variation phenomena are due to the overall communication setting, others are due to diastratic factors such as gender, provenance, age, and social background. The analysis of these factors, although relevant for both linguistic and speech technology communities, is hampered by the need to annotate existing corpora or to recruit, categorise, and record volunteers as a function of targeted profiles. This paper presents a methodology that uses a knowledge base to provide speaker-specific information. This can facilitate the enrichment of existing corpora with new annotations extracted from the knowledge base. The method also helps the large scale analysis by automatically extracting instances of speech variation to correlate with diastratic features. We apply our method to an over 120-hour corpus of broadcast speech in French and investigate variation patterns linked to reduction phenomena and/or specific to connected speech such as disfluencies. We find significant differences in speech rate, the use of filler words, and the rate of non-canonical realisations of frequent segments as a function of different professional categories and age groups. Yaru Wu, Fabian M. Suchanek, Ioana Vasilescu, Lori Lamel, Martine Adda-Decker |
LREC | 3 |
| 2021 | Synchronic Fortition in Five Romance Languages? A Large Corpus-Based Study of Word-Initial DevoicingabstractInternational audience Mathilde Hutin, Yaru Wu, Adèle Jatteau, Ioana Vasilescu, Lori Lamel, Martine Adda-Decker |
Interspeech | 4 |
| 2020 | How confident are you? Exploring the role of fillers in the automatic prediction of a speaker's confidenceabstract"Fillers", example "um" in English, have been linked to the "Feeling of Another’s Knowing (FOAK)" or the listener’s perception of a speaker’s expressed confidence. Yet, in Spoken Language Processing (SLP) they remain unexplored, or overlooked as noise. We introduce a new and challenging task, that is the prediction of FOAK, which we think has widespread applicability, given the increasing popularity of automatic processing of educational and job interviews, reviews and speeches. We design a set of filler features based on linguistic literature, and investigate their potential in FOAK prediction. We show that the integration of information related to implicature meanings allows an improvement in the FOAK model and that the different functions of fillers are differently correlated with confidence. Tanvi Dinkar, Ioana Vasilescu, Catherine Pelachaud, Chloé Clavel |
ICASSP | 2 |
| 2020 | Ongoing Phonologization of Word-Final Voicing Alternations in Two Romance Languages: Romanian and FrenchabstractInternational audience Mathilde Hutin, Adèle Jatteau, Ioana Vasilescu, Lori Lamel, Martine Adda-Decker |
INTERSPEECH | 3 |
| 2019 | " Gra[f] e!" Word-Final Devoicing of Obstruents in Standard French: An Acoustic Study Based on Large CorporaabstractInternational audience Adèle Jatteau, Ioana Vasilescu, Lori Lamel, Martine Adda-Decker, Nicolas Audibert |
INTERSPEECH | 2 |
| 2018 | Exploring Temporal Reduction in Dialectal Spanish: A Large-scale Study of Lenition of Voiced Stops and Coda-sabstractInternational audience Ioana Vasilescu, Nidia Hernández, Bianca Vieru-Dimulescu, Lori Lamel |
INTERSPEECH | 1 |
| 2016 | Marginal Contrast Among Romanian Vowels: Evidence from ASR and Functional LoadabstractInternational audience Margaret E. L. Renwick, Ioana Vasilescu, Camille Dutrey, Lori Lamel, Bianca Vieru-Dimulescu |
INTERSPEECH | 2 |
| 2014 | A CRF-based approach to automatic disfluency detection in a French call-centre corpusabstractInternational audience Camille Dutrey, Chloé Clavel, Sophie Rosset, Ioana Vasilescu, Martine Adda-Decker |
INTERSPEECH | 4 |
| 2014 | Morpho-Syntactic Study of Errors from Speech Recognition System
Maria Goryainova, Cyril Grouin, Sophie Rosset, Ioana Vasilescu |
LREC | 4 |
| 2014 | Human annotation of ASR error regions: Is "gravity" a sharable concept for human annotators?
Daniel Luzzati, Cyril Grouin, Ioana Vasilescu, Martine Adda-Decker, Eric Bilinski, Nathalie Camelin, Juliette Kahn, Carole Lailler, Lori Lamel, Sophie Rosset |
LREC | 3 |
| 2012 | Cross-lingual studies of ASR errors: paradigms for perceptual evaluations
Ioana Vasilescu, Martine Adda-Decker, Lori Lamel |
LREC | 1 |
| 2011 | Cross-Lingual Study of ASR Errors: On the Role of the Context in Human Perception of Near-HomophonesabstractInternational audience Ioana Vasilescu, Dahbia Yahia, Natalie D. Snoeren, Martine Adda-Decker, Lori Lamel |
INTERSPEECH | 1 |
| 2011 | Fiction support for realistic portrayals of fear-type emotional manifestations
Chloé Clavel, Ioana Vasilescu, Laurence Devillers |
Comput. Speech Lang. | 2 |
| 2010 | On the Role of Discourse Markers in Interactive Spoken Question Answering Systems
Ioana Vasilescu, Sophie Rosset, Martine Adda-Decker |
LREC | 1 |
| 2009 | A perceptual investigation of speech transcription errors involving frequent near-homophones in French and american EnglishabstractThis article compares the errors made by automatic speech recognizers to those made by humans for near-homophones in American English and French. This exploratory study focuses on the impact of limited word context and the potential resulting ambiguities for automatic speech recognition (ASR) systems and human listeners. Perceptual experiments using 7-gram chunks centered on incorrect or correct words output by an ASR system, show that humans make significantly more transcription errors on the first type of stimuli, thus highlighting the local ambiguity. The long-term aim of this study is to improve the modeling of such ambiguous items in order to reduce ASR errors. Ioana Vasilescu, Martine Adda-Decker, Lori Lamel, Pierre A. Hallé |
INTERSPEECH | 1 |
| 2008 | Speech Errors on Frequently Observed Homophones in French: Perceptual Evaluation vs Automatic Classification
Rena Nemoto, Ioana Vasilescu, Martine Adda-Decker |
LREC | 2 |
| 2008 | Fear-type emotion recognition for future audio-based surveillance systems
Chloé Clavel, Ioana Vasilescu, Laurence Devillers, Gaël Richard, Thibaut Ehrette |
Speech Commun. | 2 |
| 2007 | Detection and Analysis of Abnormal Situations Through Fear-Type Acoustic ManifestationsabstractRecent work on emotional speech processing has demonstrated the interest to consider the information conveyed by the emotional component in speech to enhance the understanding of human behaviors. But to date, there has been little integration of emotion detection systems in effective applications. The present research focuses on the development of a fear-type emotions recognition system to detect and analyze abnormal situations for surveillance applications. The Fear vs. Neutral classification gets a mean accuracy rate at 70.3%. It corresponds to quite optimistic results given the diversity of fear manifestations illustrated in the data. More specific acoustic models are built inside the fear class by considering the context of emergence of the emotional manifestations, i.e. the type of the threat during which they occur, and which has a strong influence on fear acoustic manifestations. The potential use of these models for a threat type recognition system is also investigated. Such information about the situation can indeed be useful for surveillance systems. Chloé Clavel, Laurence Devillers, Gaël Richard, Ioana Vasilescu, Thibaut Ehrette |
ICASSP (4) | 4 |
| 2006 | Language, gender, speaking style and language proficiency as factors influencing the autonomous vocalic filler production in spontaneous speechabstractAbstract This paper deals with the factors characterizing the production of autonomous vocalic filled pauses in large spontaneous speech corpora, namely language, gender, speaking style and language proficiency. Two types of corpora are analyzed: a corpus of broadcast news in French and American English and a corpus of short talks in a conference in English spoken by native and non-native speakers. Several acoustic and prosodic parameters are evaluated and correlated with each factor, namely timbre, pitch, duration and density. Results presented here show that the timbre is correlated with language and language proficiency, whereas the duration is linked both to gender and speaking style, the latter conditioning also the hesitation density in speech. Index terms: speech disfluencies, autonomous filled pauses, L1/L2, emotional state. 1. Introduction This paper focuses on autonomous vocalic filled pauses in spontaneous speech corpora. Among the phenomena described as “disfluencies”, filled pauses represent one of the most frequently encountered across languages. Autonomous vocalic hesitations as a type of filled pause are widely represented and consist in the insertion “at any moment” in the speech flow of a lengthened vocalic segment, alone or in combination with other segments (such as a nasal coda in English). Its aim is “to announce the initiation of what is expected to be a […] delay in speaking” [1]. Autonomous vocalic hesitations occur without lexical support and are thus to be distinguished from vocal lengthening of segments belonging to lexical items (generally function words). Filled pauses have however other possible realizations, as for instance lengthened nasal consonants (“mm” in Mandarin Chinese) or demonstratives (“ano”, “eto” in Japanese) [2,3]. For the present study we consider vocalic hesitations in French (“euh”) and English (“uh”, “um” in American English; ”er” in British English). Previously autonomous vocalic hesitations have been studied in intra- and inter-language perspectives with no particular consideration of the role of the context on their acoustic and prosodic characteristics. In our former studies, we have compared autonomous vocalic hesitations in 8 languages: American English, Middle Oriental Arabic, Mandarin Chinese, French, Italian, South-American Spanish, and European Portuguese. We have focused on the support vowel of the hesitations in each considered language. The support vowel has been defined as the main vocalic segment of a hesitation, i.e. the longest and most stable realization of each item. This vowel occurs in isolation (as unique realization of the hesitation), in a diphthong or followed by a nasal consonant as in English. Among the parameters characterizing the support vowel, duration, pitch and timbre have received a particular attention. Analysis revealed that the timbre is the most language-dependent parameter characterizing vocalic hesitations. Pitch and duration help both at differentiating the hesitation vowel from vowels with similar timbre within a given language. Pitch and duration seem to show universal patterns, i.e. the main vowel of a hesitation is significantly longer than other similar intra-lexical vowels and exhibits a flat and stable F0 contour [4]. Consequently, the hypothesis has been made that timbre is a language-dependent parameter, whereas pitch and duration could be considered as language-independent features. In this study, we consider 4 factors which may play a role in the production of vocalic hesitation in spontaneous speech corpora: language; gender; spoken style and language proficiency (mother tongue vs. second language). Ioana Vasilescu, Martine Adda-Decker |
INTERSPEECH | 1 |
| 2006 | Fear-type emotions of the SAFE Corpus: annotation issues
Chloé Clavel, Ioana Vasilescu, Laurence Devillers, Thibaut Ehrette, Gaël Richard |
LREC | 2 |
| 2005 | Perceptual salience of language-specific acoustic differences in autonomous fillers across eight languagesabstractInternational audience Ioana Vasilescu, Maria Candea, Martine Adda-Decker |
INTERSPEECH | 1 |
| 2004 | Fiction database for emotion detection in abnormal situationsabstractThe present research focuses on the acquisition and annotation of vocal resources for emotion detection. We are interested in detecting emotions occurring in abnormal situations and particularly in detecting ”fear”. The present study considers a preliminary database of audiovisual sequences extracted from movie fictions. The sequences selected provide various manifestations of target emotions and are described with a multimodal annotation tool. We focus on audio cues in the annotation strategy and we use the video as support for validating the audio labels. The present article deals with the description of the methodology of data acquisition and annotation. The validation of annotation is realized via two perceptual paradigms in which the +/-video condition in stimuli presentation varies. We show the perceptual significance of the audio cues and the presence of target emotions. Ioana Vasilescu, Laurence Devillers, Chloé Clavel, Thibaut Ehrette |
INTERSPEECH | 1 |
| 2004 | Reliability of Lexical and Prosodic Cues in Two Real-life Spoken Dialog Corpora
Laurence Devillers, Ioana Vasilescu |
LREC | 2 |
| 2003 | Emotion detection in task-oriented spoken dialoguesabstractDetecting emotions in the context of automated call center services can be helpful for following the evolution of the human-computer dialogues, enabling dynamic modification of the dialogue strategies and influencing the final outcome. The emotion detection work reported here is a part of larger study aiming to model user behavior in real interactions. We make use of a corpus of real agent-client spoken dialogues in which the manifestation of emotion is quite complex, and it is common to have shaded emotions since the interlocutors attempt to control the expression of their internal attitude. Our aims are to define appropriate emotions for call center services, to annotate the dialogues and to validate the presence of emotions via perceptual tests and to find robust cues for emotion detection. In contrast to research carried out with artificial data with simulated emotions, for real-life corpora the set of appropriate emotion labels must be determined. Two studies are reported: the first investigates automatic emotion detection using linguistic information, whereas the second concerns perceptual tests for identifying emotions as well as the prosodic and textual cues which signal them. About 11% of the utterances are annotated with non-neutral emotion labels. Preliminary experiments using lexical cues detect about 70% of these labels. Laurence Devillers, Lori Lamel, Ioana Vasilescu |
ICME | 3 |
| 2003 | Prosodic cues for emotion characterization in real-life spoken dialogsabstractThis paper reports on an analysis of prosodic cues for emotion characterization in 100 natural spoken dialogs recorded at a telephone customer service center. The corpus annotated with task-dependent emotion tags which were validated by a perceptual test. Two F0 range parameters, one at the sentence level and the other at the subsegment level, emerge as the most salient cues for emotion classification. These parameters can differentiate between negative emotion (irritation/anger, anxiety/fear) and neutral attitude and confirm trends illustrated by the perceptual experiment. Laurence Devillers, Ioana Vasilescu |
INTERSPEECH | 2 |
| 2002 | Factors in human language identification
Ian Maddieson, Ioana Vasilescu |
INTERSPEECH | 2 |
| 2001 | From perceptual designs to linguistic typology and automatic language identification : overview and perspectivesabstractThis paper deals with the overview of the methods in perceptual language identification and the suggestion of a new approach based on a two-step methodology integrating to perception “genetic” considerations and resulting into the modeling of perceptually identified discriminative cues. The first study reported here concerns experimental designs for perceptual and automatic identification of the dialectal level of languages that are less represented in the literature although it is spoken on very large geographical area (i.e. Arabic dialectal continuum). The same experimental design is implemented to determine a set of linguistic criteria for the automatic identification of 5 Romance languages (i.e. French, Italian, Spanish, Portuguese and Romanian). Melissa Barkat-Defradas, Ioana Vasilescu |
INTERSPEECH | 2 |
| 2000 | Perceptual features for the identification of Romance languagesabstractInternational audience Ioana Vasilescu, François Pellegrino, Jean-Marie Hombert |
INTERSPEECH | 1 |