EDBT 2026 Demo / reviewers in the wild / expert
Marisa Casillas
dblp:176/0420
· DBLP profile ↗
21ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0001-5417-0505ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 6 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Naturalistic observation of language development outside the home
Claire Bergey, Marisa Casillas, Daniel S. Messinger, Robert Z. Sparks |
CogSci | 2 |
| 2022 | From doggy to dog: Developmental shifts in children's use of register-specific words
Kennedy Casey, Marisa Casillas |
CogSci | 2 |
| 2022 | Sticks, leaves, buckets, and bowls: Distributional patterns of children's at-home object handling in two subsistence societies
Kennedy Casey, Mary Elliott, Elizabeth Mickiewicz, Anapaula Silva Mandujano, Kimberly Shorter, Mara Duquette, Elika Bergelson, Marisa Casillas |
CogSci | 8 |
| 2022 | Immature vocalizations simplify the speech of Tseltal Mayan and US caregivers
Steven L. Elmlinger, Michael H. Goldstein, Marisa Casillas |
CogSci | 3 |
| 2022 | Getting to the root of linguistic alignment: Testing the predictions of Interactive Alignment across developmental and biological variation in language skill
Ruthe Foushee, Daniel Byrne, Marisa Casillas, Susan Goldin-Meadow |
CogSci | 3 |
| 2021 | Analyzing contingent interactions in R with 'chattr'
Marisa Casillas, Camila Scaff |
CogSci | 1 |
| 2020 | Measuring prosodic predictability in children's home language environments
Kyle MacDonald, Marisa Casillas, Okko Johannes Räsänen, Anne S. Warlaumont |
CogSci | 2 |
| 2019 | Daylong data: Raw audio to transcript via automated \& manual open-science tools
John P. Bunce, Elika Bergelson, Anne S. Warlaumont, Marisa Casillas |
CogSci | 4 |
| 2019 | Who are you talking to like that? Exploring adults' ability to discriminate child- and adult-directed speech across languages
John P. Bunce, Melanie Soderstrom, Md Momin Al Aziz, Marisa Casillas |
CogSci | 4 |
| 2019 | The shape of language experience in two traditional communities
Marisa Casillas |
CogSci | 1 |
| 2019 | Automatic word count estimation from daylong child-centered recordings in various language environments using language-independent syllabification of speechabstractAutomatic word count estimation (WCE) from audio recordings can be used to quantify the amount of verbal communication in a recording environment. One key application of WCE is to measure language input heard by infants and toddlers in their natural environments, as captured by daylong recordings from microphones worn by the infants. Although WCE is nearly trivial for high-quality signals in high-resource languages, daylong recordings are substantially more challenging due to the unconstrained acoustic environments and the presence of near- and far-field speech. Moreover, many use cases of interest involve languages for which reliable ASR systems or even well-defined lexicons are not available. A good WCE system should also perform similarly for low- and high-resource languages in order to enable unbiased comparisons across different cultures and environments. Unfortunately, the current state-of-the-art solution, the LENA system, is based on proprietary software and has only been optimized for American English, limiting its applicability. In this paper, we build on existing work on WCE and present the steps we have taken towards a freely available system for WCE that can be adapted to different languages or dialects with a limited amount of orthographically transcribed speech data. Our system is based on language-independent syllabification of speech, followed by a language-dependent mapping from syllable counts (and a number of other acoustic features) to the corresponding word count estimates. We evaluate our system on samples from daylong infant recordings from six different corpora consisting of several languages and socioeconomic environments, all manually annotated with the same protocol to allow direct comparison. We compare a number of alternative techniques for the two key components in our system: speech activity detection and automatic syllabification of speech. As a result, we show that our system can reach relatively consistent WCE accuracy across multiple corpora and languages (with some limitations). In addition, the system outperforms LENA on three of the four corpora consisting of different varieties of English. We also demonstrate how an automatic neural network-based syllabifier, when trained on multiple languages, generalizes well to novel languages beyond the training data, outperforming two previously proposed unsupervised syllabifiers as a feature extractor for WCE. Okko Johannes Räsänen, Shreyas Seshadri, Julien Karadayi, Eric Riebling, John P. Bunce, Alejandrina Cristià, Florian Metze, Marisa Casillas, Celia Rosemberg, Elika Bergelson, Melanie Soderstrom |
Speech Commun. | 8 |
| 2018 | Talker Diarization in the Wild: the Case of Child-centered Daylong Audio-recordingsabstractSpeaker diarization (answering 'who spoke when') is a widely researched subject within speech technology. Numerous experiments have been run on datasets built from broadcast news, meeting data, and call centers—the task sometimes appears close to being solved. Much less work has begun to tackle the hardest diarization task of all: spontaneous conversations in real-world settings. Such diarization would be particularly useful for studies of language acquisition, where researchers investigate the speech children produce and hear in their daily lives. In this paper, we study audio gathered with a recorder worn by small children as they went about their normal days. As a result, each child was exposed to different acoustic environments with a multitude of background noises and a varying number of adults and peers. The inconsistency of speech and noise within and across samples poses a challenging task for speaker diarization systems, which we tackled via retraining and data augmentation techniques. We further studied sources of structured variation across raw audio files, including the impact of speaker type distribution, proportion of speech from children, and child age on diarization performance. We discuss the extent to which these findings might generalize to other samples of speech in the wild. Alejandrina Cristià, Shobhana Ganesh, Marisa Casillas, Sriram Ganapathy |
INTERSPEECH | 3 |
| 2018 | Comparison of Syllabification Algorithms and Training Strategies for Robust Word Count Estimation across Different Languages and Recording ConditionsabstractWord count estimation (WCE) from audio recordings has a number of applications, including quantifying the amount of speech that language-learning infants hear in their natural environments, as captured by daylong recordings made with devices worn by infants. To be applicable in a wide range of scenarios and also low-resource domains, WCE tools should be extremely robust against varying signal conditions and require minimal access to labeled training data in the target domain. For this purpose, earlier work has used automatic syllabification of speech, followed by a least-squares-mapping of syllables to word counts. This paper compares a number of previously proposed syllabifiers in the WCE task, including a supervised bi-directional long short-term memory (BLSTM) network that is trained on a language for which high quality syllable annotations are available (a “high resource language”), and reports how the alternative methods compare on different languages and signal conditions. We also explore additive noise and varying-channel data augmentation strategies for BLSTM training, and show how they improve performance in both matching and mismatching languages. Intriguingly, we also find that even though the BLSTM works on languages beyond its training data, the unsupervised algorithms can still outperform it in challenging signal conditions on novel languages. Okko Johannes Räsänen, Shreyas Seshadri, Marisa Casillas |
INTERSPEECH | 3 |
| 2017 | Description of the Homebank Child/Adult Addressee Corpus (HB-CHAAC)
Elika Bergelson, Andrei Amatuni, Marisa Casillas, Amanda Seidl, Melanie Soderstrom, Anne S. Warlaumont |
INTERSPEECH | 3 |
| 2017 | What do Babies Hear? Analyses of Child- and Adult-Directed SpeechabstractChild-directed speech is argued to facilitate language development, and is found cross-linguistically and cross-culturally to varying degrees. However, previous research has generally focused on short samples of child-caregiver interaction, often in the lab or with experimenters present. We test the generalizability of this phenomenon with an initial descriptive analysis of the speech heard by young children in a large, unique collection of naturalistic, daylong home recordings. Trained annotators coded automatically-detected adult speech 'utterances' from 61 homes across 4 North American cities, gathered from children (age 2-24 months) wearing audio recorders during a typical day. Coders marked the speaker gender (male/female) and intended addressee (child/adult), yielding 10,886 addressee and gender tags from 2,523 minutes of audio (cf. HB-CHAAC Interspeech ComParE challenge; Schuller et al., in press). Automated speaker-diarization (LENA) incorrectly gender-tagged 30% of male adult utterances, compared to manually-coded consensus. Furthermore, we find effects of SES and gender on child-directed and overall speech, increasing child-directed speech with child age, and interactions of speaker gender, child gender, and child age: female caretakers increased their child-directed speech more with age than male caretakers did, but only for male infants. Implications for language acquisition and existing classification algorithms are discussed. Marisa Casillas, Andrei Amatuni, Amanda Seidl, Melanie Soderstrom, Anne S. Warlaumont, Elika Bergelson |
INTERSPEECH | 1 |
| 2017 | A New Workflow for Semi-Automatized Annotations: Tests with Long-Form Naturalistic Recordings of Childrens Language EnvironmentsabstractInteroperable annotation formats are fundamental to the utility, expansion, and sustainability of collective data repositories.In language development research, shared annotation schemes have been critical to facilitating the transition from raw acoustic data to searchable, structured corpora. Current schemes typically require comprehensive and manual annotation of utterance boundaries and orthographic speech content, with an additional, optional range of tags of interest. These schemes have been enormously successful for datasets on the scale of dozens of recording hours but are untenable for long-format recording corpora, which routinely contain hundreds to thousands of audio hours. Long-format corpora would benefit greatly from (semi-)automated analyses, both on the earliest steps of annotation—voice activity detection, utterance segmentation, and speaker diarization—as well as later steps—e.g., classification-based codes such as child-vs-adult-directed speech, and speech recognition to produce phonetic/orthographic representations. We present an annotation workflow specifically designed for long-format corpora which can be tailored by individual researchers and which interfaces with the current dominant scheme for short-format recordings. The workflow allows semi-automated annotation and analyses at higher linguistic levels. We give one example of how the workflow has been successfully implemented in a large cross-database project. Marisa Casillas, Elika Bergelson, Anne S. Warlaumont, Alejandrina Cristià, Melanie Soderstrom, Mark VanDam, Han Sloetjes |
INTERSPEECH | 1 |
| 2017 | The INTERSPEECH 2017 Computational Paralinguistics Challenge: Addressee, Cold & SnoringabstractThe INTERSPEECH 2017 Computational Paralinguistics Challenge addresses three different problems for the first time in research competition under well-defined conditions: In the Addressee sub-challenge, it has to be determined whether speech produced by an adult is directed towards another adult or towards a child; in the Cold sub-challenge, speech under cold has to be told apart from ‘healthy’ speech; and in the Snoring subchallenge, four different types of snoring have to be classified. In this paper, we describe these sub-challenges, their conditions, and the baseline feature extraction and classifiers, which include data-learnt feature representations by end-to-end learning with convolutional and recurrent neural networks, and bag-of-audiowords for the first time in the challenge series Björn W. Schuller, Stefan Steidl, Anton Batliner, Elika Bergelson, Jarek Krajewski, Christoph Janott, Andrei Amatuni, Marisa Casillas, Amanda Seidl, Melanie Soderstrom, Anne S. Warlaumont, Guillermo Hidalgo, Sebastian Schnieder, Clemens Heiser, Winfried Hohenhorst, Michael Herzog, Maximilian Schmitt, Kun Qian 0003, Yue Zhang 0014, George Trigeorgis, Panagiotis Tzirakis, Stefanos Zafeiriou |
INTERSPEECH | 8 |
| 2015 | The perception of stroke-to-stroke turn boundaries in signed conversation
Marisa Casillas, Connie de Vos, Onno Crasborn, Stephen C. Levinson |
CogSci | 1 |
| 2015 | Twelve-month-olds differentiate between typical and atypical conversational timing
Elma E. Hilbrink, Marisa Casillas, Imme L. Lammertink, Stephen C. Levinson |
CogSci | 2 |
| 2013 | The development of predictive processes in children's discourse understanding
Marisa Casillas, Michael C. Frank |
CogSci | 1 |
| 2013 | Phonetic variation and the recognition of words with pronunciation variants
Meghan Sumner, Chigusa Kurumada, Roey Gafter, Marisa Casillas |
CogSci | 4 |