Alejandrina Cristià

dblp:128/2160 · DBLP profile ↗
← Back
37ranked-venue papers
2as first author
13since 2021 · last 2025
0000-0003-2979-4556ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 2 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Challenges in Automated Processing of Speech from Child Wearables: The Case of Voice Type Classifier
abstract
International audience
Tarek Kunze, Marianne Métais, Hadrien Titeux, Lucas Elbert, Joseph Coffey, Emmanuel Dupoux, Alejandrina Cristià, Marvin Lavechin
INTERSPEECH7
2025 Fifteen Years of Child-Centered Long-Form Recordings: Promises, Resources, and Remaining Challenges to Validity
abstract
Audio-recordings collected with a child-worn device are a fundamental tool in child language research. Long-form recordings collected over whole days promise to capture children's input and production with minimal observer bias, and therefore high validity. The sheer volume of resulting data necessitates automated analysis to extract relevant metrics for researchers and clinicians. This paper summarizes collective knowledge on this technique, providing entry points to existing resources. We also highlight various sources of error that threaten the accuracy of automated annotations and the interpretation of resulting metrics. To address this, we propose potential troubleshooting metrics to help users assess data quality. While a fully automated quality control system is not feasible, we outline practical strategies for researchers to improve data collection and contextualize their analyses.
Loann Peurey, Marvin Lavechin, Tarek Kunze, Manel Khentout, Lucas Gautheron, Emmanuel Dupoux, Alejandrina Cristià
INTERSPEECH7
2025 Employing self-supervised learning models for cross-linguistic child speech maturity classification
abstract
International audience
Theo Zhang, Madurya Suresh, Anne Warluamont, Kasia Hitczenko, Alejandrina Cristià, Margaret Cychosz
INTERSPEECH5
2024 The Difficulty and Importance of Estimating the Lower and Upper Bounds of Infant Speech Exposure
Joseph Coffey, Okko Johannes Räsänen, Camila Scaff, Alejandrina Cristià
INTERSPEECH4
2023 Brouhaha: Multi-Task Training for Voice Activity Detection, Speech-to-Noise Ratio, and C50 Room Acoustics Estimation
abstract
Most automatic speech processing systems register degraded performance when applied to noisy or reverberant speech. But how can one tell whether speech is noisy or reverberant? We propose Brouhaha, a neural network jointly trained to extract speech/non-speech segments, speech-to-noise ratios, and C50 room acoustics from single-channel recordings. Brouhaha is trained using a data-driven approach in which noisy and reverberant audio segments are synthesized. We first evaluate its performance and demonstrate that the proposed multi-task regime is beneficial. We then present two scenarios illustrating how Brouhaha can be used on naturally noisy and reverberant data: 1) to investigate the errors made by a speaker diarization model (pyannote.audio); and 2) to assess the reliability of an automatic speech recognition model (Whisper from OpenAI). Both our pipeline and a pretrained model are open source and shared with the speech community.
Marvin Lavechin, Marianne Métais, Hadrien Titeux, Alodie Boissonnet, Jade Copet, Morgane Rivière, Elika Bergelson, Alejandrina Cristià, Emmanuel Dupoux, Hervé Bredin
ASRU8
2023 Analysing the Impact of Audio Quality on the Use of Naturalistic Long-Form Recordings for Infant-Directed Speech Research
María Andrea Cruz Blandón, Alejandrina Cristià, Okko Johannes Räsänen
CogSci2
2023 Hearttoheart: The Arts of Infant Versus Adult-Directed Speech Classification
abstract
Psycholinguistics researchers investigate child language exposure by studying children’s language environment. A main factor is whether, in humanistic heart-to-heart dialogue, the speech is directed to the infant (infant-directed speech) versus to another adult (adult-directed speech). The former has been found to better predict children’s lexicon, and therefore constitutes a more relevant part of children’s language environment. Listening to, segmenting and annotating naturalistic long-form recordings collected through infant-worn devices is highly costly and time-consuming, and could be prone to errors in misclassification. We aim to overcome these challenges by automatically classifying speech as infant-directed versus adult-directed. In this research, we exploit multiple datasets, combined to form a larger corpus for training. In addition, we employ four different methods: Multi-task learning, adversarial training, autoencoder multi-task learning and adversarial multi-task learning, the last of which yielded the best results on all datasets.
Najla Al Futaisi, Alejandrina Cristià, Björn W. Schuller
ICASSP2
2023 〈'〉 in Tsimane': a Preliminary Investigation
abstract
Tsimane' is a language spoken in Bolivia by several thousand people.Yet, it has not been described in detail.We aim to take a step towards a better description by focusing on an aspect of language: the sound represented in spelling with 〈'〉, informally described as a glottal stop.We recorded two adult speakers of Tsimane' producing (near-)minimal pairs involving this sound.Perceptual analyses suggested 〈'〉 is very rarely realised as a full glottal stop, and is more often cued by creaky-voiced vowels and nasals.Despite the variability in implementation, presentation of syllabic minimal pairs to these two informants and two other adult Tsimane' listeners revealed evidence that they could easily perceive when 〈'〉 was intended.Together, these data suffice to rule out the hypothesis that 〈'〉 is systematically realised as a full stop, and suggests instead a more complex set of perceptual cues may be at speakers' and listeners' disposal.
William Havard, Yaya Sy, Camila Scaff, Loann Peurey, Alejandrina Cristià
INTERSPEECH5
2023 BabySLM: language-acquisition-friendly benchmark of self-supervised spoken language models
abstract
International audience
Marvin Lavechin, Yaya Sy, Hadrien Titeux, María Andrea Cruz Blandón, Okko Johannes Räsänen, Hervé Bredin, Emmanuel Dupoux, Alejandrina Cristià
INTERSPEECH8
2023 Measuring Language Development From Child-centered Recordings
abstract
Standard ways to measure child language development from spontaneous corpora rely on detailed linguistic descriptions of a language as well as exhaustive transcriptions of the child’s speech, which today can only be done through costly human labor. We tackle both issues by proposing (1) a new language development metric (based on entropy) that does not require linguistic knowledge other than having a corpus of text in the language in question to train a language model, (2) a method to derive this metric directly from speech based on a smaller text-speech parallel corpus. Here, we present descriptive results on an open archive including data from six Englishlearning children as a proof of concept. We document that our entropy metric documents a gradual convergence of children’s speech towards adults’ speech as a function of age, and it also correlates moderately with lexical and morphosyntactic measures derived from morphologically parsed transcriptions.The source code of the experiments is released at https://github.com/yaya-sy/EntropyBasedCLDMetricsIndex Terms: L1 acquisition, child speech, morphosyntax, phonetics, speech technology application
Yaya Sy, William Havard, Marvin Lavechin, Emmanuel Dupoux, Alejandrina Cristià
INTERSPEECH5
2023 Using iterative adaptation and dynamic mask for child speech extraction under real-world multilingual conditions
Shi Cheng 0001, Jun Du 0002, Shutong Niu, Alejandrina Cristià, Xin Wang 0037, Qing Wang 0008, Chin-Hui Lee 0001
Speech Commun.4
2021 Child Language Acquisition Studied with Wearables
Alejandrina Cristià
Interspeech1
2021 Towards Large-Scale Data Annotation of Audio from Wearables: Validating Zooniverse Annotations of Infant Vocalization Types
abstract
Recent developments allow the collection of audio data from lightweight wearable devices, potentially enabling us to study language use from everyday life samples. However, extracting useful information from these data is currently impossible with automatized routines, and overly expensive with trained human annotators. We explore a strategy fit to the 21st century, relying on the collaboration of citizen scientists. A large dataset of infant speech was uploaded on a citizen science platform. The same data were annotated in the laboratory by highly trained annotators. We investigate whether crowd-sourced annotations are qualitatively and quantitatively com-parable to those produced by expert annotators in a dataset of children at high- and low-risk for language disorders. Our results reveal that classification of individual vocalizations on Zooniverse was overall moderately accurate compared to the laboratory gold standard. The analysis of descriptors defined at the level of individual children found strong correlations between descriptors derived from Zooniverse versus laboratory annotations.
Chiara Semenzin, Lisa Hamrick, Amanda Seidl, Bridgette Kelleher, Alejandrina Cristià
SLT5
2020 A Study of Child Speech Extraction Using Joint Speech Enhancement and Separation in Realistic Conditions
abstract
In this paper, we design a novel joint framework of speech enhancement and speech separation for child speech extraction in realistic conditions, targeting the problem of extracting child speech from daily conversations in BabyTrain mega corpus. To the best of our knowledge, it is the first discussion of a feasible method for child speech extraction in realistic conditions. First, we make detailed analysis of the BabyTrain mega corpus, which is recorded in adverse environments. We observe problems of background noises, reverberations and child speech that is partially obscured by adult speech (for instance due to speaker overlap but also imitation by the adult). Motivated by this, we conduct a joint framework of speech enhancement and speech separation for child speech extraction. To measure the extraction results in realistic conditions, we propose several objective measurements to evaluate the performance of the our system, which is different from those commonly used for simulation data. Compared with the unprocessed approach and classification approach, our proposed approach can yield the best performance among all subsets of BabyTrain.
Xin Wang 0037, Jun Du 0002, Alejandrina Cristià, Lei Sun 0010, Chin-Hui Lee 0001
ICASSP3
2020 An Open-Source Voice Type Classifier for Child-Centered Daylong Recordings
abstract
International audience
Marvin Lavechin, Ruben Bousbib, Hervé Bredin, Emmanuel Dupoux, Alejandrina Cristià
INTERSPEECH5
2020 Seshat: a Tool for Managing and Verifying Annotation Campaigns of Audio Data
abstract
We introduce Seshat, a new, simple and open-source software to efficiently manage annotations of speech corpora. The Seshat software allows users to easily customise and manage annotations of large audio corpora while ensuring compliance with the formatting and naming conventions of the annotated output files. In addition, it includes procedures for checking the content of annotations following specific rules that can be implemented in personalised parsers. Finally, we propose a double-annotation mode, for which Seshat computes automatically an associated inter-annotator agreement with the gamma measure taking into account the categorisation and segmentation discrepancies.
Hadrien Titeux, Rachid Riad, Xuan-Nga Cao, Nicolas Hamilakis, Kris Madden, Alejandrina Cristià, Anne-Catherine Bachoud-Lévi, Emmanuel Dupoux
LREC6
2019 Is Word Segmentation Child's Play in All Languages?
abstract
When learning language, infants need to break down the flow of input speech into minimal word-like units, a process best described as unsupervised bottom-up segmentation.Proposed strategies include several segmentation algorithms, but only cross-linguistically robust algorithms could be plausible candidates for human word learning, since infants have no initial knowledge of the ambient language.We report on the stability in performance of 11 conceptually diverse algorithms on a selection of 8 typologically distinct languages.The results are evidence that some segmentation algorithms are cross-linguistically valid, thus could be considered as potential strategies employed by all infants.
Georgia-Rengina Loukatou, Steven Moran, Damián E. Blasi, Sabine Stoll, Alejandrina Cristià
ACL (1)5
2019 Is it easier to segment words from infant- than adult-directed speech? Modeling evidence from an ecological French corpus
Georgia-Rengina Loukatou, Marie-Thérèse Le Normand, Alejandrina Cristià
CogSci3
2019 VCMNet: Weakly Supervised Learning for Automatic Infant Vocalisation Maturity Analysis
abstract
Using neural networks to classify infant vocalisations into important subclasses (such as crying versus speech) is an emergent task in speech technology. One of the biggest roadblocks standing in the way of progress lies in the datasets: The performance of a learning model is affected by the labelling quality and size of the dataset used, and infant vocalisation datasets with good quality labels tend to be small. In this paper, we assess the performance of three models for infant VoCalisation Maturity (VCM) trained with a large dataset annotated automatically using a purpose-built classifier and a small dataset annotated by highly trained human coders. The two datasets are used in three different training strategies, whose performance is compared against a baseline model. The first training strategy investigates adversarial training, while the second exploits multi-task learning as the neural network trains on both datasets simultaneously. In the final strategy, we integrate adversarial training and multi-task learning. All of the training strategies outperform the baseline, with the adversarial training strategy yielding the best results on the development set.
Najla Al Futaisi, Zixing Zhang 0001, Alejandrina Cristià, Anne S. Warlaumont, Björn W. Schuller
ICMI3
2019 The Second DIHARD Diarization Challenge: Dataset, Task, and Baselines
abstract
This paper introduces the second DIHARD challenge, the second in a series of speaker diarization challenges intended to improve the robustness of diarization systems to variation in recording equipment, noise conditions, and conversational domain. The challenge comprises four tracks evaluating diarization performance under two input conditions (single channel vs. multi-channel) and two segmentation conditions (diarization from a reference speech segmentation vs. diarization from scratch). In order to prevent participants from overtuning to a particular combination of recording conditions and conversational domain, recordings are drawn from a variety of sources ranging from read audiobooks to meeting speech, to child language acquisition recordings, to dinner parties, to web video. We describe the task and metrics, challenge design, datasets, and baseline systems for speech enhancement, speech activity detection, and diarization.
Neville Ryant, Kenneth Church 0001, Christopher Cieri, Alejandrina Cristià, Jun Du 0002, Sriram Ganapathy, Mark Y. Liberman
INTERSPEECH4
2019 The INTERSPEECH 2019 Computational Paralinguistics Challenge: Styrian Dialects, Continuous Sleepiness, Baby Sounds & Orca Activity
abstract
The INTERSPEECH 2019 Computational Paralinguistics Challenge addresses four different problems for the first time in a research competition under well-defined conditions: In the Styrian Dialects Sub-Challenge, three types of Austrian-German dialects have to be classified; in the Continuous Sleepiness Sub-Challenge, the sleepiness of a speaker has to be assessed as regression problem; in the Baby Sound Sub-Challenge, five types of infant sounds have to be classified; and in the Orca Activity Sub-Challenge, orca sounds have to be detected.We describe the Sub-Challenges and baseline feature extraction and classifiers, which include data-learnt (supervised) feature representations by the 'usual' ComParE and BoAW features, and deep unsupervised representation learning using the AUDEEP toolkit.
Björn W. Schuller, Anton Batliner, Christian Bergler, Florian B. Pokorny, Jarek Krajewski, Margaret Cychosz, Ralf Vollmann, Sonja-Dana Roelen, Sebastian Schnieder, Elika Bergelson, Alejandrina Cristià, Amanda Seidl, Anne S. Warlaumont, Lisa Yankowitz, Elmar Nöth, Shahin Amiriparian, Simone Hantke, Maximilian Schmitt
INTERSPEECH11
2019 Towards Detection of Canonical Babbling by Citizen Scientists: Performance as a Function of Clip Length
Amanda Seidl, Anne S. Warlaumont, Alejandrina Cristià
INTERSPEECH3
2019 Automatic word count estimation from daylong child-centered recordings in various language environments using language-independent syllabification of speech
abstract
Automatic word count estimation (WCE) from audio recordings can be used to quantify the amount of verbal communication in a recording environment. One key application of WCE is to measure language input heard by infants and toddlers in their natural environments, as captured by daylong recordings from microphones worn by the infants. Although WCE is nearly trivial for high-quality signals in high-resource languages, daylong recordings are substantially more challenging due to the unconstrained acoustic environments and the presence of near- and far-field speech. Moreover, many use cases of interest involve languages for which reliable ASR systems or even well-defined lexicons are not available. A good WCE system should also perform similarly for low- and high-resource languages in order to enable unbiased comparisons across different cultures and environments. Unfortunately, the current state-of-the-art solution, the LENA system, is based on proprietary software and has only been optimized for American English, limiting its applicability. In this paper, we build on existing work on WCE and present the steps we have taken towards a freely available system for WCE that can be adapted to different languages or dialects with a limited amount of orthographically transcribed speech data. Our system is based on language-independent syllabification of speech, followed by a language-dependent mapping from syllable counts (and a number of other acoustic features) to the corresponding word count estimates. We evaluate our system on samples from daylong infant recordings from six different corpora consisting of several languages and socioeconomic environments, all manually annotated with the same protocol to allow direct comparison. We compare a number of alternative techniques for the two key components in our system: speech activity detection and automatic syllabification of speech. As a result, we show that our system can reach relatively consistent WCE accuracy across multiple corpora and languages (with some limitations). In addition, the system outperforms LENA on three of the four corpora consisting of different varieties of English. We also demonstrate how an automatic neural network-based syllabifier, when trained on multiple languages, generalizes well to novel languages beyond the training data, outperforming two previously proposed unsupervised syllabifiers as a feature extractor for WCE.
Okko Johannes Räsänen, Shreyas Seshadri, Julien Karadayi, Eric Riebling, John P. Bunce, Alejandrina Cristià, Florian Metze, Marisa Casillas, Celia Rosemberg, Elika Bergelson, Melanie Soderstrom
Speech Commun.6
2018 Enhancement and Analysis of Conversational Speech: JSALT 2017
abstract
Automatic speech recognition is more and more widely and effectively used. Nevertheless, in some automatic speech analysis tasks the state of the art is surprisingly poor. One of these is “diarization”, the task of determining who spoke when. Diarization is key to processing meeting audio and clinical interviews, extended recordings such as police body cam or child language acquisition data, and any other speech data involving multiple speakers whose voices are not cleanly separated into individual channels. Overlapping speech, environmental noise and suboptimal recording techniques make the problem harder. During the JSALT Summer Workshop at CMU in 2017, an international team of researchers worked on several aspects of this problem, including calibration of the state of the art, detection of overlaps, enhancement of noisy recordings, and classification of shorter speech segments. This paper sketches the workshop's results, and announces plans for a “Diarization Challenge” to encourage further progress.
Neville Ryant, Elika Bergelson, Kenneth Church 0001, Alejandrina Cristià, Jun Du 0002, Sriram Ganapathy, Sanjeev Khudanpur, Diana Kowalski, Mahesh Krishnamoorthy, Rajat Kulshreshta, Mark Y. Liberman, Yu-Ding Lu, Matthew Maciejewski, Florian Metze, Ján Profant, Lei Sun 0010, Yu Tsao 0001
ICASSP4
2018 Talker Diarization in the Wild: the Case of Child-centered Daylong Audio-recordings
abstract
Speaker diarization (answering 'who spoke when') is a widely researched subject within speech technology. Numerous experiments have been run on datasets built from broadcast news, meeting data, and call centers—the task sometimes appears close to being solved. Much less work has begun to tackle the hardest diarization task of all: spontaneous conversations in real-world settings. Such diarization would be particularly useful for studies of language acquisition, where researchers investigate the speech children produce and hear in their daily lives. In this paper, we study audio gathered with a recorder worn by small children as they went about their normal days. As a result, each child was exposed to different acoustic environments with a multitude of background noises and a varying number of adults and peers. The inconsistency of speech and noise within and across samples poses a challenging task for speaker diarization systems, which we tackled via retraining and data augmentation techniques. We further studied sources of structured variation across raw audio files, including the impact of speaker type distribution, proportion of speech from children, and child age on diarization performance. We discuss the extent to which these findings might generalize to other samples of speech in the wild.
Alejandrina Cristià, Shobhana Ganesh, Marisa Casillas, Sriram Ganapathy
INTERSPEECH1
2018 The ACLEW DiViMe: An Easy-to-use Diarization Tool
Adrien Le Franc, Eric Riebling, Julien Karadayi, Yun Wang 0005, Camila Scaff, Florian Metze, Alejandrina Cristià
INTERSPEECH7
2018 Automated Classification of Children's Linguistic versus Non-Linguistic Vocalisations
abstract
A key outstanding task for speech technology involves dealing with non-standard speakers, notably young children.Distinguishing children's linguistic from non-linguistic vocalisations is crucial for a number of applied and fundamental research goals, and yet there are few systems available for such a classification.This paper investigates two large-scale framelevel acoustic feature sets (eGeMAPS and ComParE16) followed by a dynamic model (GRU-RNN), and two kinds of derived static feature sets on the segment level (functional-based and Bag of Audio Words) combined with a static model (SVM), and automatically learnt representations directly from original raw voice signals by using an end-to-end system.These are applied to a large database of children's vocalisations (total N = 6,298) drawn from daylong recordings gathered in Namibia, Bolivia, and Vanuatu.Among these systems, the one implemented with GRU-RNN using ComParE16 features empirically performs best.We further identify promising paths of further research, including the application of a finer-grained classification of children's vocalisations onto these data, and the exploration of other feature systems.
Zixing Zhang 0001, Alejandrina Cristià, Anne S. Warlaumont, Björn W. Schuller
INTERSPEECH2
2017 Top-Down versus Bottom-Up Theories of Phonological Acquisition: A Big Data Approach
abstract
Recent work has made available a number of standardized meta-analyses bearing on various aspects of infant language processing. We utilize data from two such meta-analyses (discrimination of vowel contrasts and word segmentation, i.e., recognition of word forms extracted from running speech) to assess whether the published body of empirical evidence supports a bottom-up versus a top-down theory of early phonological development by leveling the power of results from thousands of infants. We predicted that if infants can rely purely on auditory experience to develop their phonological categories, then vowel discrimination and word segmentation should develop in parallel, with the latter being potentially lagged compared to the former. However, if infants crucially rely on word form information to build their phonological categories, then development at the word level must precede the acquisition of native sound categories. Our results do not support the latter prediction. We discuss potential implications and limitations, most saliently that word forms are only one topdown level proposed to affect phonological development, with other proposals suggesting that top-down pressures emerge from lexical (i.e., word-meaning pairs) development. This investigation also highlights general procedures by which standardized meta-analyses may be reused to answer theoretical questions spanning across phenomena.
Christina Bergmann, Sho Tsuji, Alejandrina Cristià
INTERSPEECH3
2017 A New Workflow for Semi-Automatized Annotations: Tests with Long-Form Naturalistic Recordings of Childrens Language Environments
abstract
Interoperable annotation formats are fundamental to the utility, expansion, and sustainability of collective data repositories.In language development research, shared annotation schemes have been critical to facilitating the transition from raw acoustic data to searchable, structured corpora. Current schemes typically require comprehensive and manual annotation of utterance boundaries and orthographic speech content, with an additional, optional range of tags of interest. These schemes have been enormously successful for datasets on the scale of dozens of recording hours but are untenable for long-format recording corpora, which routinely contain hundreds to thousands of audio hours. Long-format corpora would benefit greatly from (semi-)automated analyses, both on the earliest steps of annotation—voice activity detection, utterance segmentation, and speaker diarization—as well as later steps—e.g., classification-based codes such as child-vs-adult-directed speech, and speech recognition to produce phonetic/orthographic representations. We present an annotation workflow specifically designed for long-format corpora which can be tailored by individual researchers and which interfaces with the current dominant scheme for short-format recordings. The workflow allows semi-automated annotation and analyses at higher linguistic levels. We give one example of how the workflow has been successfully implemented in a large cross-database project.
Marisa Casillas, Elika Bergelson, Anne S. Warlaumont, Alejandrina Cristià, Melanie Soderstrom, Mark VanDam, Han Sloetjes
INTERSPEECH4
2017 Relating Unsupervised Word Segmentation to Reported Vocabulary Acquisition
abstract
International audience
Elin Larsen, Alejandrina Cristià, Emmanuel Dupoux
INTERSPEECH2
2017 MetaLab: A Repository for Meta-Analyses on Language Development, and More
Sho Tsuji, Christina Bergmann, Molly Lewis, Mika Braginsky, Page Piccinini, Michael C. Frank, Alejandrina Cristià
INTERSPEECH7
2017 Which Acoustic and Phonological Factors Shape Infants' Vowel Discrimination? Exploiting Natural Variation in InPhonDB
abstract
A key research question in early language acquisition concerns\n\nthe development of infants’ ability to discriminate sounds, and\n\nthe factors structuring discrimination abilities. Vowel discrimination,\n\nin particular, has been studied using a range of tasks, experimental\n\nparadigms, and stimuli over the past 40 years, work\n\nrecently compiled in a meta-analysis. We use this meta-analysis\n\nto assess whether there is statistical evidence for the following\n\nfactors affecting effect sizes across studies: (1) the order in\n\nwhich the two vowel stimuli are presented; and (2) the distance\n\nbetween the vowels, measured acoustically in terms of spectral\n\nand quantity differences. The magnitude of effect sizes analysis\n\nrevealed order effects consistent with the Natural Referent\n\nVowels framework, with greater effect sizes when the second\n\nvowel was more peripheral than the first. Additionally, we find\n\nthat spectral acoustic distinctiveness is a consistent predictor of\n\nstudies’ effect sizes, while temporal distinctiveness did not predict\n\neffect size magnitude. None of these factors interacted significantly\n\nwith age. We discuss implications of these results for\n\nlanguage acquisition, and more generally developmental psychology,\n\nresearch.\n\nA key research question in early language acquisition concerns\n\nthe development of infants’ ability to discriminate sounds, and\n\nthe factors structuring discrimination abilities. Vowel discrimination,\n\nin particular, has been studied using a range of tasks, experimental\n\nparadigms, and stimuli over the past 40 years, work\n\nrecently compiled in a meta-analysis. We use this meta-analysis\n\nto assess whether there is statistical evidence for the following\n\nfactors affecting effect sizes across studies: (1) the order in\n\nwhich the two vowel stimuli are presented; and (2) the distance\n\nbetween the vowels, measured acoustically in terms of spectral\n\nand quantity differences. The magnitude of effect sizes analysis\n\nrevealed order effects consistent with the Natural Referent\n\nVowels framework, with greater effect sizes when the second\n\nvowel was more peripheral than the first. Additionally, we find\n\nthat spectral acoustic distinctiveness is a consistent predictor of\n\nstudies’ effect sizes, while temporal distinctiveness did not predict\n\neffect size magnitude. None of these factors interacted significantly\n\nwith age. We discuss implications of these results for\n\nlanguage acquisition, and more generally developmental psychology,\n\nresearch.
Sho Tsuji, Alejandrina Cristià
INTERSPEECH2
2017 HomeBank: A Repository for Long-Form Real-World Audio Recordings of Children
Anne S. Warlaumont, Mark VanDam, Elika Bergelson, Alejandrina Cristià
INTERSPEECH4
2016 Discriminability of sound contrasts in the face of speaker variation quantified
Christina Bergmann, Alejandrina Cristià, Emmanuel Dupoux
CogSci2
2016 Tutorial: Meta-Analytic Methods for Cognitive Science
Sho Tsuji, Molly Lewis, Christina Bergmann, Mike Frank, Alejandrina Cristià
CogSci5
2015 Salient dimensions in implicit phonotactic learning
abstract
Adults are able to learn sound co-occurrences without conscious knowledge after brief exposures. But which dimensions of sounds are most salient in this process? Using an artificial phonology paradigm, we explored potential learnability differences involving consonant-, speaker-, and tone-vowel cooccurrences. Results revealed that participants, whose native language was not tonal, implicitly encoded consonant-vowel patterns with a high level of accuracy; were above chance for tone-vowel co-occurrences; and were at chance for speakervowel co-occurrences. This pattern of results is exactly what would be expected if both language-specific experience and innate biases to encode potentially contrastive linguistic dimensions affect the salience of different dimensions during implicit learning of sound patterns.
Elise Michon, Emmanuel Dupoux, Alejandrina Cristià
INTERSPEECH3
2014 Acoustic correlates of phonological status
abstract
Languages vary not only in terms of their sound inventory, but also in the phonological status certain sound distinctions are assigned. For example, while vowel nasality is lexically contrastive (phonemic) in Quebecois French, it is largely determined by the context (allophonic) in American English; the reverse is true for vowel tenseness. If phonetics and phonology interact, a minimal pair of sounds should span a larger acoustic divergence when it is pronounced by speakers for whom the underlying distinction is phonemic compared to allophonic. Near minimal pairs were segmented from a corpus of American English and Quebecois French using a crossed design (since nasality and tenseness have opposite phonological status in the two languages). Pairwise time-aligned divergences between contrasts were calculated on the basis of 7 mainstream spoken feature representations, and a set of linguistic phonetic measurements. Only carefully selected phonetic measurements revealed the expected cross-over, with larger divergences for English than French tokens of the tenseness contrast, and larger divergences for French than English tokens for the nasality contrast. We conclude that the phonetic effects of phonological status are subtle enough that only linguistically-informed (or supervised) measurements can pick up on them.
Maarten Versteegh, Amanda Seidl, Alejandrina Cristià
INTERSPEECH3