Abdellah Fourtassi

dblp:136/5079 · DBLP profile ↗
← Back
29ranked-venue papers
8as first author
20since 2021 · last 2025
0000-0003-0279-7730ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 8 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 7 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 Benchmarking LLMs for Mimicking Child-Caregiver Language in Interaction
Abdellah Fourtassi
CogSci2
2024 Analysing Communicative Intent Coordination in Child-Caregiver Interactions
Abhishek Agrawal, Benoît Favre, Abdellah Fourtassi
CogSci3
2024 Child-Caregiver Gaze Dynamics in Naturalistic Face-to-Face Conversations
Dhia-Elhak Goumri, Leonor Becerra-Bonache, Abdellah Fourtassi
CogSci3
2024 Development of Flexible Role-Taking in Conversations Across Preschool
Morgane Peirolo, Abdellah Fourtassi
CogSci3
2024 Automatic Coding of Contingency in Child-Caregiver Conversations
abstract
One of the most important communicative skills children have to learn is to engage in meaningful conversations with people around them. At the heart of this learning lies the mastery of contingency, i.e., the ability to contribute to an ongoing exchange in a relevant fashion (e.g., by staying on topic). Current research on this question relies on the manual annotation of a small sample of children, which limits our ability to draw general conclusions about development. Here, we propose to mitigate the limitations of manual labor by relying on automatic tools for contingency judgment in children’s early natural interactions with caregivers. Drawing inspiration from the field of dialogue systems evaluation, we built and compared several automatic classifiers. We found that a Transformer-based pre-trained language model – when fine-tuned on a relatively small set of data we annotated manually (around 3,500 turns) – provided the best predictions. We used this model to automatically annotate, new and large-scale data, almost two orders of magnitude larger than our fine-tuning set. It was able to replicate existing results and generate new data-driven hypotheses. The broad impact of the work is to provide resources that can help the language development community study communicative development at scale, leading to more robust theories.
Abhishek Agrawal, Mitja Nikolaus, Benoît Favre, Abdellah Fourtassi
LREC/COLING4
2024 CHICA: A Developmental Corpus of Child-Caregiver's Face-to-face vs. Video Call Conversations in Middle Childhood
abstract
Existing studies of naturally occurring language-in-interaction have largely focused on the two ends of the developmental spectrum, i.e., early childhood and adulthood, leaving a gap in our knowledge about how development unfolds, especially across middle childhood. The current work contributes to filling this gap by introducing CHICA (for Child Interpersonal Communication Analysis), a developmental corpus of child-caregiver conversations at home, involving groups of French-speaking children aged 7, 9, and 11 years old. Each dyad was recorded twice: once in a face-to-face setting and once using computer-mediated video calls. For the face-to-face settings, we capitalized on recent advances in mobile, lightweight eye-tracking and head motion detection technology to optimize the naturalness of the recordings, allowing us to obtain both precise and ecologically valid data. Further, we mitigated the challenges of manual annotation by relying – to the extent possible – on automatic tools in speech processing and computer vision. Finally, to demonstrate the richness of this corpus for the study of child communicative development, we provide preliminary analyses comparing several measures of child-caregiver conversational dynamics across developmental age, modality, and communicative medium. We hope the current corpus will allow new discoveries into the properties and mechanisms of multimodal communicative development across middle childhood.
Dhia-Elhak Goumri, Abhishek Agrawal, Mitja Nikolaus, Hong Duc Thang Vu, Kübra Bodur, Elias Emmar, Cassandre Armand, Chiara Mazzocconi, Shreejata Gupta, Laurent Prévot 0001, Benoît Favre, Leonor Becerra-Bonache, Abdellah Fourtassi
LREC/COLING13
2024 Automatic Annotation of Grammaticality in Child-Caregiver Conversations
abstract
The acquisition of grammar has been a central question to adjudicate between theories of language acquisition. In order to conduct faster, more reproducible, and larger-scale corpus studies on grammaticality in child-caregiver conversations, tools for automatic annotation can offer an effective alternative to tedious manual annotation. We propose a coding scheme for context-dependent grammaticality in child-caregiver conversations and annotate more than 4,000 utterances from a large corpus of transcribed conversations. Based on these annotations, we train and evaluate a range of NLP models. Our results show that fine-tuned Transformer-based models perform best, achieving human inter-annotation agreement levels. As a first application and sanity check of this tool, we use the trained models to annotate a corpus almost two orders of magnitude larger than the manually annotated data and verify that children’s grammaticality shows a steady increase with age. This work contributes to the growing literature on applying state-of-the-art NLP methods to help study child language acquisition at scale.
Mitja Nikolaus, Abhishek Agrawal, Petros Kaklamanis, Alex Warstadt, Abdellah Fourtassi
LREC/COLING5
2024 Language Learning, Representation, and Processing in Humans and Machines: Introduction to the Special Issue
abstract
Abstract Large Language Models (LLMs) and humans acquire knowledge about language without direct supervision. LLMs do so by means of specific training objectives, while humans rely on sensory experience and social interaction. This parallelism has created a feeling in NLP and cognitive science that a systematic understanding of how LLMs acquire and use the encoded knowledge could provide useful insights for studying human cognition. Conversely, methods and findings from the field of cognitive science have occasionally inspired language model development. Yet, the differences in the way that language is processed by machines and humans—in terms of learning mechanisms, amounts of data used, grounding and access to different modalities—make a direct translation of insights challenging. The aim of this edited volume has been to create a forum of exchange and debate along this line of research, inviting contributions that further elucidate similarities and differences between humans and LLMs.
Marianna Apidianaki, Abdellah Fourtassi, Sebastian Padó
Comput. Linguistics2
2023 Development of Multimodal Turn Coordination in Conversations: Evidence for Adult-like behavior in Middle Childhood
Abhishek Agrawal, Kübra Bodur, Benoît Favre, Abdellah Fourtassi
CogSci5
2023 Comparing Children and Large Language Models in Word Sense Disambiguation: Insights and Challenges
Francesco Cabiddu, Mitja Nikolaus, Abdellah Fourtassi
CogSci3
2023 Differences between Mimicking and Non-Mimicking laughter in Child-Caregiver Conversation: A Distributional and Acoustic Analysis
Chiara Mazzocconi, Benjamin O'Brien, Kevin El Haddad, Kübra Bodur, Abdellah Fourtassi
CogSci5
2023 Typology of topological relations using machine translation
Floor Meewis, Abdellah Fourtassi, Isabelle Dautriche
CogSci2
2023 Communicative Feedback in Response to Children's Grammatical Errors
Mitja Nikolaus, Laurent Prévot 0001, Abdellah Fourtassi
CogSci3
2022 Backchannel Behavior in Child-Caregiver Zoom-Mediated Conversations
Kübra Bodur, Mitja Nikolaus, Abdellah Fourtassi, Laurent Prévot 0001
CogSci3
2022 Communicative Feedback as a Mechanism Supporting the Production of Intelligible Speech in Early Childhood
Mitja Nikolaus, Laurent Prévot 0001, Abdellah Fourtassi
CogSci3
2022 Do Vision-and-Language Transformers Learn Grounded Predicate-Noun Dependencies?
abstract
Recent advances in vision-and-language modeling have seen the development of Transformer architectures that achieve remarkable performance on multimodal reasoning tasks.Yet, the exact capabilities of these black-box models are still poorly understood.While much of previous work has focused on studying their ability to learn meaning at the word-level, their ability to track syntactic dependencies between words has received less attention.We take a first step in closing this gap by creating a new multimodal task targeted at evaluating understanding of predicate-noun dependencies in a controlled setup.We evaluate a range of state-of-the-art models and find that their performance on the task varies considerably, with some models performing relatively well and others at chance level.In an effort to explain this variability, our analyses indicate that the quality (and not only sheer quantity) of pretraining data is essential.Additionally, the best performing models leverage fine-grained multimodal pretraining objectives in addition to the standard image-text matching objectives.This study highlights that targeted and controlled evaluations are a crucial step for a precise and rigorous test of the multimodal knowledge of vision-and-language models.
Mitja Nikolaus, Emmanuelle Salin, Stéphane Ayache, Abdellah Fourtassi, Benoît Favre
EMNLP4
2021 On the Role of Low-level Linguistic Levels for Reading Time Prediction
Franck Dary, Abdellah Fourtassi, Alexis Nasr
CogSci2
2021 Large-scale study of speech acts' development using automatic labelling
Mitja Nikolaus, Juliette Maes, Jérémy Auguste, Laurent Prévot 0001, Abdellah Fourtassi
CogSci5
2021 Modeling speech act development in early childhood: the role of frequency and linguistic cues
Mitja Nikolaus, Juliette Maes, Abdellah Fourtassi
CogSci3
2021 Modeling the Interaction Between Perception-Based and Production-Based Learning in Children's Early Acquisition of Semantic Knowledge
abstract
Children learn the meaning of words and sentences in their native language at an impressive speed and from highly ambiguous input.To account for this learning, previous computational modeling has focused mainly on the study of perception-based mechanisms like cross-situational learning.However, children do not learn only by exposure to the input.As soon as they start to talk, they practice their knowledge in social interactions and they receive feedback from their caregivers.In this work, we propose a model integrating both perception-and production-based learning using artificial neural networks which we train on a large corpus of crowd-sourced images with corresponding descriptions.We found that production-based learning improves performance above and beyond perception-based learning across a wide range of semantic tasks including both word-and sentence-level semantics.In addition, we documented a synergy between these two mechanisms, where their alternation allows the model to converge on more balanced semantic knowledge.The broader impact of this work is to highlight the importance of modeling language learning in the context of social interactions where children are not only understood as passively absorbing the input, but also as actively participating in the construction of their linguistic knowledge.
Mitja Nikolaus, Abdellah Fourtassi
CoNLL2
2020 Discovering Conceptual Hierarchy Through Explicit and Implicit Cues in Child-Directed Speech
Abdellah Fourtassi, Kyra Wilson, Michael C. Frank
CogSci1
2019 Phoneme learning is influenced by the taxonomic similarity of the semantic referents
Abdellah Fourtassi, Emmanuel Dupoux
CogSci1
2019 Continuous developmental change can explain discontinuities in word learning
Abdellah Fourtassi, Sophie Regan, Michael C. Frank
CogSci1
2018 Word Learning as Network Growth: A Cross-linguistic Analysis
Abdellah Fourtassi, Yuan Bian 0004, Michael C. Frank
CogSci1
2017 Word Identification Under Multimodal Uncertainty
Abdellah Fourtassi, Michael C. Frank
CogSci1
2016 The role of word-word co-occurrence in word learning
Abdellah Fourtassi, Emmanuel Dupoux
CogSci1
2014 Self-Consistency as an Inductive Bias in Early Language Acquisition
Abdellah Fourtassi, Ewan Dunbar, Emmanuel Dupoux
CogSci1
2014 A Rudimentary Lexicon and Semantics Help Bootstrap Phoneme Acquisition
abstract
Infants spontaneously discover the relevant phonemes of their language without any direct supervision.This acquisition is puzzling because it seems to require the availability of high levels of linguistic structures (lexicon, semantics), that logically suppose the infants having a set of phonemes already.We show how this circularity can be broken by testing, in realsize language corpora, a scenario whereby infants would learn approximate representations at all levels, and then refine them in a mutually constraining way.We start with corpora of spontaneous speech that have been encoded in a varying number of detailed context-dependent allophones.We derive, in an unsupervised way, an approximate lexicon and a rudimentary semantic representation.Despite the fact that all these representations are poor approximations of the ground truth, they help reorganize the fine grained categories into phoneme-like categories with a high degree of accuracy.
Abdellah Fourtassi, Emmanuel Dupoux
CoNLL1
2013 A summary of the 2012 JHU CLSP workshop on zero resource speech technologies and models of early language acquisition
abstract
We summarize the accomplishments of a multi-disciplinary workshop exploring the computational and scientific issues surrounding zero resource (unsupervised) speech technologies and related models of early language acquisition. Centered around the tasks of phonetic and lexical discovery, we consider unified evaluation metrics, present two new approaches for improving speaker independence in the absence of supervision, and evaluate the application of Bayesian word segmentation algorithms to automatic subword unit tokenizations. Finally, we present two strategies for integrating zero resource techniques into supervised settings, demonstrating the potential of unsupervised methods to improve mainstream technologies.
Aren Jansen, Emmanuel Dupoux, Sharon Goldwater, Mark Johnson 0001, Sanjeev Khudanpur, Kenneth Church 0001, Naomi Feldman, Hynek Hermansky, Florian Metze, Richard C. Rose, Mike Seltzer, Pascal Clark, Ian McGraw, Balakrishnan Varadarajan, Erin D. Bennett, Benjamin Börschinger, Justin T. Chiu, Ewan Dunbar, Abdellah Fourtassi, David F. Harwath, Chia-ying Lee, Keith D. Levin, Atta Norouzian, Vijayaditya Peddinti, Rachael Richardson, Thomas Schatz, Samuel Thomas 0001
ICASSP19