Afra Alishahi

dblp:69/8699 · DBLP profile ↗
← Back
30ranked-venue papers
5as first author
10since 2021 · last 2026
0009-0009-9237-3253ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 5 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Gender Disambiguation in Machine Translation: Diagnostic Evaluation in Decoder-Only Architectures
abstract
While Large Language Models achieve state-of-the-art results across a wide range of NLP tasks, they remain prone to systematic biases. Among these, gender bias is particularly salient in MT, due to systematic differences across languages in whether and how gender is marked. As a result, translation often requires disambiguating implicit source signals into explicit gender-marked forms. In this context, standard benchmarks may capture broad disparities but fail to reflect the full complexity of gender bias in modern MT. In this paper, we extend recent frameworks on bias evaluation by: (i) introducing a novel measure coined ’Prior Bias’, capturing a model’s default gender assumptions, and (ii) applying the framework to decoder-only MT models. Our results show that, despite their scale and state-of-the-art status, decoder-only models do not generally outperform encoder-decoder architectures on gender-specific metrics; however, post-training (e.g., instruction tuning) not only improves contextual awareness but also reduces the masculine Prior Bias.
Chiara Manna, Hosein Mohebbi, Afra Alishahi, Frédéric Blain, Eva Vanmassenhove
LREC3
2026 Mechanistic Interpretability Meets Cognitive Linguistics: Modelling Locative Image Schemas in the Circuit Framework
Mattia Proietti, Afra Alishahi, Grzegorz Chrupala, Alessandro Lenci
LREC2
2025 On the reliability of feature attribution methods for speech classification
abstract
As the capabilities of large-scale pre-trained models evolve, understanding the determinants of their outputs becomes more important. Feature attribution aims to reveal which parts of the input elements contribute the most to model outputs. In speech processing, the unique characteristics of the input signal make the application of feature attribution methods challenging. We study how factors such as input type and aggregation and perturbation timespan impact the reliability of standard feature attribution methods, and how these factors interact with characteristics of each classification task. We find that standard approaches to feature attribution are generally unreliable when applied to the speech domain, with the exception of word-aligned perturbation methods when applied to word-based classification tasks.
Gaofei Shen, Hosein Mohebbi, Arianna Bisazza, Afra Alishahi, Grzegorz Chrupala
INTERSPEECH4
2024 Encoding of lexical tone in self-supervised models of spoken language
abstract
Gaofei Shen, Michaela Watkins, Afra Alishahi, Arianna Bisazza, Grzegorz Chrupała. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Gaofei Shen, Michaela Watkins, Afra Alishahi, Arianna Bisazza, Grzegorz Chrupala
NAACL-HLT3
2024 Perception of Phonological Assimilation by Neural Speech Recognition Models
abstract
Abstract Human listeners effortlessly compensate for phonological changes during speech perception, often unconsciously inferring the intended sounds. For example, listeners infer the underlying /n/ when hearing an utterance such as “clea[m] pan”, where [m] arises from place assimilation to the following labial [p]. This article explores how the neural speech recognition model Wav2Vec2 perceives assimilated sounds, and identifies the linguistic knowledge that is implemented by the model to compensate for assimilation during Automatic Speech Recognition (ASR). Using psycholinguistic stimuli, we systematically analyze how various linguistic context cues influence compensation patterns in the model’s output. Complementing these behavioral experiments, our probing experiments indicate that the model shifts its interpretation of assimilated sounds from their acoustic form to their underlying form in its final layers. Finally, our causal intervention experiments suggest that the model relies on minimal phonological context cues to accomplish this shift. These findings represent a step towards better understanding the similarities and differences in phonological processing between neural ASR models and humans.
Charlotte Pouw, Marianne de Heer Kloots, Afra Alishahi, Willem H. Zuidema
Comput. Linguistics3
2023 Quantifying Context Mixing in Transformers
abstract
Self-attention weights and their transformed variants have been the main source of information for analyzing token-to-token interactions in Transformer-based models. But despite their ease of interpretation, these weights are not faithful to the models’ decisions as they are only one part of an encoder, and other components in the encoder layer can have considerable impact on information mixing in the output representations. In this work, by expanding the scope of analysis to the whole encoder block, we propose Value Zeroing, a novel context mixing score customized for Transformers that provides us with a deeper understanding of how information is mixed at each encoder layer. We demonstrate the superiority of our context mixing score over other analysis methods through a series of complementary evaluations with different viewpoints based on linguistically informed rationales, probing, and faithfulness analysis.
Hosein Mohebbi, Willem H. Zuidema, Grzegorz Chrupala, Afra Alishahi
EACL4
2023 Homophone Disambiguation Reveals Patterns of Context Mixing in Speech Transformers
abstract
Transformers have become a key architecture in speech processing, but our understanding of how they build up representations of acoustic and linguistic structure is limited.In this study, we address this gap by investigating how measures of 'context-mixing' developed for text models can be adapted and applied to models of spoken language.We identify a linguistic phenomenon that is ideal for such a case study: homophony in French (e.g.livre vs livres), where a speech recognition model has to attend to syntactic cues such as determiners and pronouns in order to disambiguate spoken words with identical pronunciations and transcribe them while respecting grammatical agreement.We perform a series of controlled experiments and probing analyses on Transformer-based speech models.Our findings reveal that representations in encoder-only models effectively incorporate these cues to identify the correct transcription, whereas encoders in encoder-decoder models mainly relegate the task of capturing contextual dependencies to decoder modules. 1
Hosein Mohebbi, Grzegorz Chrupala, Willem H. Zuidema, Afra Alishahi
EMNLP4
2023 Linguistic Productivity: the Case of Determiners in English
abstract
Raquel G. Alhama, Ruthe Foushee, Daniel Byrne, Allyson Ettinger, Susan Goldin-Meadow, Afra Alishahi. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Raquel G. Alhama, Ruthe Foushee, Daniel Byrne, Allyson Ettinger, Susan Goldin-Meadow, Afra Alishahi
IJCNLP (1)6
2023 Wave to Syntax: Probing spoken language models for syntax
abstract
Understanding which information is encoded in deep models of spoken and written language has been the focus of much research in recent years, as it is crucial for debugging and improving these architectures. Most previous work has focused on probing for speaker characteristics, acoustic and phonological information in models of spoken language, and for syntactic information in models of written language. Here we focus on the encoding of syntax in several self-supervised and visually grounded models of spoken language. We employ two complementary probing methods, combined with baselines and reference representations to quantify the degree to which syntactic structure is encoded in the activations of the target models. We show that syntax is captured most prominently in the middle layers of the networks, and more explicitly within models with more parameters.
Gaofei Shen, Afra Alishahi, Arianna Bisazza, Grzegorz Chrupala
INTERSPEECH2
2022 Learning English with Peppa Pig
abstract
Abstract Recent computational models of the acquisition of spoken language via grounding in perception exploit associations between spoken and visual modalities and learn to represent speech and visual data in a joint vector space. A major unresolved issue from the point of ecological validity is the training data, typically consisting of images or videos paired with spoken descriptions of what is depicted. Such a setup guarantees an unrealistically strong correlation between speech and the visual data. In the real world the coupling between the linguistic and the visual modality is loose, and often confounded by correlations with non-semantic aspects of the speech signal. Here we address this shortcoming by using a dataset based on the children’s cartoon Peppa Pig. We train a simple bi-modal architecture on the portion of the data consisting of dialog between characters, and evaluate on segments containing descriptive narrations. Despite the weak and confounded signal in this training data, our model succeeds at learning aspects of the visual semantics of spoken language.
Mitja Nikolaus, Afra Alishahi, Grzegorz Chrupala
Trans. Assoc. Comput. Linguistics2
2020 Analyzing analytical methods: The case of phonology in neural models of spoken language
abstract
Given the fast development of analysis techniques for NLP and speech processing systems, few systematic studies have been conducted to compare the strengths and weaknesses of each method. As a step in this direction we study the case of representations of phonology in neural network models of spoken language. We use two commonly applied analytical techniques, diagnostic classifiers and representational similarity analysis, to quantify to what extent neural activation patterns encode phonemes and phoneme sequences. We manipulate two factors that can affect the outcome of analysis. First, we investigate the role of learning by comparing neural activations extracted from trained versus randomly-initialized models. Second, we examine the temporal scope of the activations by probing both local activations corresponding to a few milliseconds of the speech signal, and global activations pooled over the whole utterance. We conclude that reporting analysis results with randomly initialized models is crucial, and that global-scope methods tend to yield more consistent results and we recommend their use as a complement to local-scope diagnostic methods.
Grzegorz Chrupala, Bertrand Higy, Afra Alishahi
ACL3
2020 Learning to Understand Child-directed and Adult-directed Speech
abstract
Speech directed to children differs from adult-directed speech in linguistic aspects such as repetition, word choice, and sentence length, as well as in aspects of the speech signal itself, such as prosodic and phonemic variation. Human language acquisition research indicates that child-directed speech helps language learners. This study explores the effect of child-directed speech when learning to extract semantic information from speech directly. We compare the task performance of models trained on adult-directed speech (ADS) and child-directed speech (CDS). We find indications that CDS helps in the initial stages of learning, but eventually, models trained on ADS reach comparable task performance, and generalize better. The results suggest that this is at least partially due to linguistic rather than acoustic properties of the two registers, as we see the same pattern when looking at models trained on acoustically comparable synthetic speech.
Lieke Gelderloos, Grzegorz Chrupala, Afra Alishahi
ACL3
2020 Active Word Learning through Self-supervision
Lieke Gelderloos, Alireza Mahmoudi Kamelabad, Afra Alishahi
CogSci3
2019 Correlating Neural and Symbolic Representations of Language
abstract
Analysis methods which enable us to better understand the representations and functioning of neural models of language are increasingly needed as deep learning becomes the dominant approach in NLP.Here we present two methods based on Representational Similarity Analysis (RSA) and Tree Kernels (TK) which allow us to directly quantify how strongly the information encoded in neural activation patterns corresponds to information represented by symbolic structures such as syntax trees.We first validate our methods on the case of a simple synthetic language for arithmetic expressions with clearly defined syntax and semantics, and show that they exhibit the expected pattern of results.We then apply our methods to correlate neural representations of English sentences with their constituency parse trees.
Grzegorz Chrupala, Afra Alishahi
ACL (1)2
2019 Curious Topics: A Curiosity-Based Model of First Language Word Learning
Daan Keijser, Lieke Gelderloos, Afra Alishahi
CogSci3
2019 Analyzing and interpreting neural networks for NLP: A report on the first BlackboxNLP workshop
abstract
Abstract The Empirical Methods in Natural Language Processing (EMNLP) 2018 workshop BlackboxNLP was dedicated to resources and techniques specifically developed for analyzing and understanding the inner-workings and representations acquired by neural models of language. Approaches included: systematic manipulation of input to neural networks and investigating the impact on their performance, testing whether interpretable knowledge can be decoded from intermediate representations acquired by neural networks, proposing modifications to neural network architectures to make their knowledge state or generated output more explainable, and examining the performance of networks on simplified or formal languages. Here we review a number of representative studies in each category.
Afra Alishahi, Grzegorz Chrupala, Tal Linzen
Nat. Lang. Eng.1
2018 Revisiting the Hierarchical Multiscale LSTM
abstract
Hierarchical Multiscale LSTM (Chung et. al., 2016) is a state-of-the-art language model that learns interpretable structure from character-level input. Such models can provide fertile ground for (cognitive) computational linguistics studies. However, the high complexity of the architecture, training and implementations might hinder its applicability. We provide a detailed reproduction and ablation study of the architecture, shedding light on some of the potential caveats of re-purposing complex deep-learning architectures. We further show that simplifying certain aspects of the architecture can in fact improve its performance. We also investigate the linguistic units (segments) learned by various levels of the model, and argue that their quality does not correlate with the overall performance of the model on language modeling.
Ákos Kádár, Marc-Alexandre Côté, Grzegorz Chrupala, Afra Alishahi
COLING4
2018 Lessons Learned in Multilingual Grounded Language Learning
abstract
Recent work has shown how to learn better visual-semantic embeddings by leveraging image descriptions in more than one language.Here, we investigate in detail which conditions affect the performance of this type of grounded language learning model.We show that multilingual training improves over bilingual training, and that low-resource languages benefit from training with higher-resource languages.We demonstrate that a multilingual model can be trained equally well on either translations or comparable sentence pairs, and that annotating the same set of images in multiple language enables further improvements via an additional caption-caption ranking objective.
Ákos Kádár, Desmond Elliott, Marc-Alexandre Côté, Grzegorz Chrupala, Afra Alishahi
CoNLL5
2017 Representations of language in a model of visually grounded speech signal
abstract
We present a visually grounded model of speech perception which projects spoken utterances and images to a joint semantic space.We use a multi-layer recurrent highway network to model the temporal nature of spoken speech, and show that it learns to extract both form and meaningbased linguistic knowledge from the input signal.We carry out an in-depth analysis of the representations used by different components of the trained model and show that encoding of semantic aspects tends to become richer as we go up the hierarchy of layers, whereas encoding of formrelated aspects of the language input tends to initially increase and then plateau or decrease.
Grzegorz Chrupala, Lieke Gelderloos, Afra Alishahi
ACL (1)3
2017 Encoding of phonology in a recurrent neural model of grounded speech
abstract
We study the representation and encoding of phonemes in a recurrent neural network model of grounded speech.We use a model which processes images and their spoken descriptions, and projects the visual and auditory representations into the same semantic space.We perform a number of analyses on how information about individual phonemes is encoded in the MFCC features extracted from the speech signal, and the activations of the layers of the model.Via experiments with phoneme decoding and phoneme discrimination we show that phoneme representations are most salient in the lower layers of the model, where low-level signals are processed at a fine-grained level, although a large amount of phonological information is retain at the top recurrent layer.We further find out that the attention mechanism following the top recurrent layer significantly attenuates encoding of phonology and makes the utterance embeddings much more invariant to synonymy.Moreover, a hierarchical clustering of phoneme representations learned by the network shows an organizational structure of phonemes similar to those proposed in linguistics.
Afra Alishahi, Marie Barking, Grzegorz Chrupala
CoNLL1
2017 Representation of Linguistic Form and Function in Recurrent Neural Networks
abstract
We present novel methods for analyzing the activation patterns of recurrent neural networks from a linguistic point of view and explore the types of linguistic structure they learn. As a case study, we use a standard standalone language model, and a multi-task gated recurrent network architecture consisting of two parallel pathways with shared word embeddings: The Visual pathway is trained on predicting the representations of the visual scene corresponding to an input sentence, and the Textual pathway is trained to predict the next word in the same sentence. We propose a method for estimating the amount of contribution of individual tokens in the input to the final prediction of the networks. Using this method, we show that the Visual pathway pays selective attention to lexical categories and grammatical functions that carry semantic information, and learns to treat word types differently depending on their grammatical function and their position in the sequential structure of the sentence. In contrast, the language models are comparatively more sensitive to words with a syntactic function. Further analysis of the most informative n-gram contexts for each model shows that in comparison with the Visual pathway, the language models react more strongly to abstract contexts that represent syntactic constructions.
Ákos Kádár, Grzegorz Chrupala, Afra Alishahi
Comput. Linguistics3
2016 A connectionist model for automatic generation of child-adult interaction patterns
Moinuddin M. Haque, Paul Vogt, Afra Alishahi, Emiel Krahmer
CogSci3
2015 Distributional determinants of learning argument structure constructions in first and second language
Yevgen Matusevych, Afra Alishahi, Ad Backus
CogSci2
2014 The impact of emerging knowledge of linguistic structure on word learning
Eva van den Bemd, Afra Alishahi, Maria Mos
CogSci2
2014 Isolating second language learning factors in a computational study of bilingual construction acquisition
Yevgen Matusevych, Afra Alishahi, Ad Backus
CogSci2
2013 Automatic generation of naturalistic child-adult interaction data
Yevgen Matusevych, Afra Alishahi, Paul Vogt
CogSci2
2012 Concurrent Acquisition of Word Meaning and Lexical Categories
Afra Alishahi, Grzegorz Chrupala
EMNLP-CoNLL1
2011 The Onset of Syntactic Bootstrapping in Word Learning: Evidence from a Computational Study
Afra Alishahi, Pirita Pyykkönen
CogSci1
2010 Online Entropy-Based Model of Lexical Category Acquisition
Grzegorz Chrupala, Afra Alishahi
CoNLL2
2008 Fast Mapping in Word Learning: What Probabilities Tell Us
Afra Alishahi, Afsaneh Fazly, Suzanne Stevenson
CoNLL1