VLDB 2026 Research / reviewers in the wild / expert
Afra Alishahi
dblp:69/8699
· DBLP profile ↗
30ranked-venue papers
5as first author
10since 2021 · last 2026
0009-0009-9237-3253ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 5 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Gender Disambiguation in Machine Translation: Diagnostic Evaluation in Decoder-Only ArchitecturesabstractWhile Large Language Models achieve state-of-the-art results across a wide range of NLP tasks, they remain prone to systematic biases. Among these, gender bias is particularly salient in MT, due to systematic differences across languages in whether and how gender is marked. As a result, translation often requires disambiguating implicit source signals into explicit gender-marked forms. In this context, standard benchmarks may capture broad disparities but fail to reflect the full complexity of gender bias in modern MT. In this paper, we extend recent frameworks on bias evaluation by: (i) introducing a novel measure coined ’Prior Bias’, capturing a model’s default gender assumptions, and (ii) applying the framework to decoder-only MT models. Our results show that, despite their scale and state-of-the-art status, decoder-only models do not generally outperform encoder-decoder architectures on gender-specific metrics; however, post-training (e.g., instruction tuning) not only improves contextual awareness but also reduces the masculine Prior Bias. Chiara Manna, Hosein Mohebbi, Afra Alishahi, Frédéric Blain, Eva Vanmassenhove |
LREC | 3 |
| 2026 | Mechanistic Interpretability Meets Cognitive Linguistics: Modelling Locative Image Schemas in the Circuit Framework
Mattia Proietti, Afra Alishahi, Grzegorz Chrupala, Alessandro Lenci |
LREC | 2 |
| 2025 | On the reliability of feature attribution methods for speech classificationabstractAs the capabilities of large-scale pre-trained models evolve, understanding the determinants of their outputs becomes more important. Feature attribution aims to reveal which parts of the input elements contribute the most to model outputs. In speech processing, the unique characteristics of the input signal make the application of feature attribution methods challenging. We study how factors such as input type and aggregation and perturbation timespan impact the reliability of standard feature attribution methods, and how these factors interact with characteristics of each classification task. We find that standard approaches to feature attribution are generally unreliable when applied to the speech domain, with the exception of word-aligned perturbation methods when applied to word-based classification tasks. Gaofei Shen, Hosein Mohebbi, Arianna Bisazza, Afra Alishahi, Grzegorz Chrupala |
INTERSPEECH | 4 |
| 2024 | Encoding of lexical tone in self-supervised models of spoken languageabstractGaofei Shen, Michaela Watkins, Afra Alishahi, Arianna Bisazza, Grzegorz Chrupała. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Gaofei Shen, Michaela Watkins, Afra Alishahi, Arianna Bisazza, Grzegorz Chrupala |
NAACL-HLT | 3 |
| 2024 | Perception of Phonological Assimilation by Neural Speech Recognition ModelsabstractAbstract Human listeners effortlessly compensate for phonological changes during speech perception, often unconsciously inferring the intended sounds. For example, listeners infer the underlying /n/ when hearing an utterance such as “clea[m] pan”, where [m] arises from place assimilation to the following labial [p]. This article explores how the neural speech recognition model Wav2Vec2 perceives assimilated sounds, and identifies the linguistic knowledge that is implemented by the model to compensate for assimilation during Automatic Speech Recognition (ASR). Using psycholinguistic stimuli, we systematically analyze how various linguistic context cues influence compensation patterns in the model’s output. Complementing these behavioral experiments, our probing experiments indicate that the model shifts its interpretation of assimilated sounds from their acoustic form to their underlying form in its final layers. Finally, our causal intervention experiments suggest that the model relies on minimal phonological context cues to accomplish this shift. These findings represent a step towards better understanding the similarities and differences in phonological processing between neural ASR models and humans. Charlotte Pouw, Marianne de Heer Kloots, Afra Alishahi, Willem H. Zuidema |
Comput. Linguistics | 3 |
| 2023 | Quantifying Context Mixing in TransformersabstractSelf-attention weights and their transformed variants have been the main source of information for analyzing token-to-token interactions in Transformer-based models. But despite their ease of interpretation, these weights are not faithful to the models’ decisions as they are only one part of an encoder, and other components in the encoder layer can have considerable impact on information mixing in the output representations. In this work, by expanding the scope of analysis to the whole encoder block, we propose Value Zeroing, a novel context mixing score customized for Transformers that provides us with a deeper understanding of how information is mixed at each encoder layer. We demonstrate the superiority of our context mixing score over other analysis methods through a series of complementary evaluations with different viewpoints based on linguistically informed rationales, probing, and faithfulness analysis. Hosein Mohebbi, Willem H. Zuidema, Grzegorz Chrupala, Afra Alishahi |
EACL | 4 |
| 2023 | Homophone Disambiguation Reveals Patterns of Context Mixing in Speech TransformersabstractTransformers have become a key architecture in speech processing, but our understanding of how they build up representations of acoustic and linguistic structure is limited.In this study, we address this gap by investigating how measures of 'context-mixing' developed for text models can be adapted and applied to models of spoken language.We identify a linguistic phenomenon that is ideal for such a case study: homophony in French (e.g.livre vs livres), where a speech recognition model has to attend to syntactic cues such as determiners and pronouns in order to disambiguate spoken words with identical pronunciations and transcribe them while respecting grammatical agreement.We perform a series of controlled experiments and probing analyses on Transformer-based speech models.Our findings reveal that representations in encoder-only models effectively incorporate these cues to identify the correct transcription, whereas encoders in encoder-decoder models mainly relegate the task of capturing contextual dependencies to decoder modules. 1 Hosein Mohebbi, Grzegorz Chrupala, Willem H. Zuidema, Afra Alishahi |
EMNLP | 4 |
| 2023 | Linguistic Productivity: the Case of Determiners in EnglishabstractRaquel G. Alhama, Ruthe Foushee, Daniel Byrne, Allyson Ettinger, Susan Goldin-Meadow, Afra Alishahi. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Raquel G. Alhama, Ruthe Foushee, Daniel Byrne, Allyson Ettinger, Susan Goldin-Meadow, Afra Alishahi |
IJCNLP (1) | 6 |
| 2023 | Wave to Syntax: Probing spoken language models for syntaxabstractUnderstanding which information is encoded in deep models of spoken and written language has been the focus of much research in recent years, as it is crucial for debugging and improving these architectures. Most previous work has focused on probing for speaker characteristics, acoustic and phonological information in models of spoken language, and for syntactic information in models of written language. Here we focus on the encoding of syntax in several self-supervised and visually grounded models of spoken language. We employ two complementary probing methods, combined with baselines and reference representations to quantify the degree to which syntactic structure is encoded in the activations of the target models. We show that syntax is captured most prominently in the middle layers of the networks, and more explicitly within models with more parameters. Gaofei Shen, Afra Alishahi, Arianna Bisazza, Grzegorz Chrupala |
INTERSPEECH | 2 |
| 2022 | Learning English with Peppa PigabstractAbstract Recent computational models of the acquisition of spoken language via grounding in perception exploit associations between spoken and visual modalities and learn to represent speech and visual data in a joint vector space. A major unresolved issue from the point of ecological validity is the training data, typically consisting of images or videos paired with spoken descriptions of what is depicted. Such a setup guarantees an unrealistically strong correlation between speech and the visual data. In the real world the coupling between the linguistic and the visual modality is loose, and often confounded by correlations with non-semantic aspects of the speech signal. Here we address this shortcoming by using a dataset based on the children’s cartoon Peppa Pig. We train a simple bi-modal architecture on the portion of the data consisting of dialog between characters, and evaluate on segments containing descriptive narrations. Despite the weak and confounded signal in this training data, our model succeeds at learning aspects of the visual semantics of spoken language. Mitja Nikolaus, Afra Alishahi, Grzegorz Chrupala |
Trans. Assoc. Comput. Linguistics | 2 |
| 2020 | Analyzing analytical methods: The case of phonology in neural models of spoken languageabstractGiven the fast development of analysis techniques for NLP and speech processing systems, few systematic studies have been conducted to compare the strengths and weaknesses of each method. As a step in this direction we study the case of representations of phonology in neural network models of spoken language. We use two commonly applied analytical techniques, diagnostic classifiers and representational similarity analysis, to quantify to what extent neural activation patterns encode phonemes and phoneme sequences. We manipulate two factors that can affect the outcome of analysis. First, we investigate the role of learning by comparing neural activations extracted from trained versus randomly-initialized models. Second, we examine the temporal scope of the activations by probing both local activations corresponding to a few milliseconds of the speech signal, and global activations pooled over the whole utterance. We conclude that reporting analysis results with randomly initialized models is crucial, and that global-scope methods tend to yield more consistent results and we recommend their use as a complement to local-scope diagnostic methods. Grzegorz Chrupala, Bertrand Higy, Afra Alishahi |
ACL | 3 |
| 2020 | Learning to Understand Child-directed and Adult-directed SpeechabstractSpeech directed to children differs from adult-directed speech in linguistic aspects such as repetition, word choice, and sentence length, as well as in aspects of the speech signal itself, such as prosodic and phonemic variation. Human language acquisition research indicates that child-directed speech helps language learners. This study explores the effect of child-directed speech when learning to extract semantic information from speech directly. We compare the task performance of models trained on adult-directed speech (ADS) and child-directed speech (CDS). We find indications that CDS helps in the initial stages of learning, but eventually, models trained on ADS reach comparable task performance, and generalize better. The results suggest that this is at least partially due to linguistic rather than acoustic properties of the two registers, as we see the same pattern when looking at models trained on acoustically comparable synthetic speech. Lieke Gelderloos, Grzegorz Chrupala, Afra Alishahi |
ACL | 3 |
| 2020 | Active Word Learning through Self-supervision
Lieke Gelderloos, Alireza Mahmoudi Kamelabad, Afra Alishahi |
CogSci | 3 |
| 2019 | Correlating Neural and Symbolic Representations of LanguageabstractAnalysis methods which enable us to better understand the representations and functioning of neural models of language are increasingly needed as deep learning becomes the dominant approach in NLP.Here we present two methods based on Representational Similarity Analysis (RSA) and Tree Kernels (TK) which allow us to directly quantify how strongly the information encoded in neural activation patterns corresponds to information represented by symbolic structures such as syntax trees.We first validate our methods on the case of a simple synthetic language for arithmetic expressions with clearly defined syntax and semantics, and show that they exhibit the expected pattern of results.We then apply our methods to correlate neural representations of English sentences with their constituency parse trees. Grzegorz Chrupala, Afra Alishahi |
ACL (1) | 2 |
| 2019 | Curious Topics: A Curiosity-Based Model of First Language Word Learning
Daan Keijser, Lieke Gelderloos, Afra Alishahi |
CogSci | 3 |
| 2019 | Analyzing and interpreting neural networks for NLP: A report on the first BlackboxNLP workshopabstractAbstract The Empirical Methods in Natural Language Processing (EMNLP) 2018 workshop BlackboxNLP was dedicated to resources and techniques specifically developed for analyzing and understanding the inner-workings and representations acquired by neural models of language. Approaches included: systematic manipulation of input to neural networks and investigating the impact on their performance, testing whether interpretable knowledge can be decoded from intermediate representations acquired by neural networks, proposing modifications to neural network architectures to make their knowledge state or generated output more explainable, and examining the performance of networks on simplified or formal languages. Here we review a number of representative studies in each category. Afra Alishahi, Grzegorz Chrupala, Tal Linzen |
Nat. Lang. Eng. | 1 |
| 2018 | Revisiting the Hierarchical Multiscale LSTMabstractHierarchical Multiscale LSTM (Chung et. al., 2016) is a state-of-the-art language model that learns interpretable structure from character-level input. Such models can provide fertile ground for (cognitive) computational linguistics studies. However, the high complexity of the architecture, training and implementations might hinder its applicability. We provide a detailed reproduction and ablation study of the architecture, shedding light on some of the potential caveats of re-purposing complex deep-learning architectures. We further show that simplifying certain aspects of the architecture can in fact improve its performance. We also investigate the linguistic units (segments) learned by various levels of the model, and argue that their quality does not correlate with the overall performance of the model on language modeling. Ákos Kádár, Marc-Alexandre Côté, Grzegorz Chrupala, Afra Alishahi |
COLING | 4 |
| 2018 | Lessons Learned in Multilingual Grounded Language LearningabstractRecent work has shown how to learn better visual-semantic embeddings by leveraging image descriptions in more than one language.Here, we investigate in detail which conditions affect the performance of this type of grounded language learning model.We show that multilingual training improves over bilingual training, and that low-resource languages benefit from training with higher-resource languages.We demonstrate that a multilingual model can be trained equally well on either translations or comparable sentence pairs, and that annotating the same set of images in multiple language enables further improvements via an additional caption-caption ranking objective. Ákos Kádár, Desmond Elliott, Marc-Alexandre Côté, Grzegorz Chrupala, Afra Alishahi |
CoNLL | 5 |
| 2017 | Representations of language in a model of visually grounded speech signalabstractWe present a visually grounded model of speech perception which projects spoken utterances and images to a joint semantic space.We use a multi-layer recurrent highway network to model the temporal nature of spoken speech, and show that it learns to extract both form and meaningbased linguistic knowledge from the input signal.We carry out an in-depth analysis of the representations used by different components of the trained model and show that encoding of semantic aspects tends to become richer as we go up the hierarchy of layers, whereas encoding of formrelated aspects of the language input tends to initially increase and then plateau or decrease. Grzegorz Chrupala, Lieke Gelderloos, Afra Alishahi |
ACL (1) | 3 |
| 2017 | Encoding of phonology in a recurrent neural model of grounded speechabstractWe study the representation and encoding of phonemes in a recurrent neural network model of grounded speech.We use a model which processes images and their spoken descriptions, and projects the visual and auditory representations into the same semantic space.We perform a number of analyses on how information about individual phonemes is encoded in the MFCC features extracted from the speech signal, and the activations of the layers of the model.Via experiments with phoneme decoding and phoneme discrimination we show that phoneme representations are most salient in the lower layers of the model, where low-level signals are processed at a fine-grained level, although a large amount of phonological information is retain at the top recurrent layer.We further find out that the attention mechanism following the top recurrent layer significantly attenuates encoding of phonology and makes the utterance embeddings much more invariant to synonymy.Moreover, a hierarchical clustering of phoneme representations learned by the network shows an organizational structure of phonemes similar to those proposed in linguistics. Afra Alishahi, Marie Barking, Grzegorz Chrupala |
CoNLL | 1 |
| 2017 | Representation of Linguistic Form and Function in Recurrent Neural NetworksabstractWe present novel methods for analyzing the activation patterns of recurrent neural networks from a linguistic point of view and explore the types of linguistic structure they learn. As a case study, we use a standard standalone language model, and a multi-task gated recurrent network architecture consisting of two parallel pathways with shared word embeddings: The Visual pathway is trained on predicting the representations of the visual scene corresponding to an input sentence, and the Textual pathway is trained to predict the next word in the same sentence. We propose a method for estimating the amount of contribution of individual tokens in the input to the final prediction of the networks. Using this method, we show that the Visual pathway pays selective attention to lexical categories and grammatical functions that carry semantic information, and learns to treat word types differently depending on their grammatical function and their position in the sequential structure of the sentence. In contrast, the language models are comparatively more sensitive to words with a syntactic function. Further analysis of the most informative n-gram contexts for each model shows that in comparison with the Visual pathway, the language models react more strongly to abstract contexts that represent syntactic constructions. Ákos Kádár, Grzegorz Chrupala, Afra Alishahi |
Comput. Linguistics | 3 |
| 2016 | A connectionist model for automatic generation of child-adult interaction patterns
Moinuddin M. Haque, Paul Vogt, Afra Alishahi, Emiel Krahmer |
CogSci | 3 |
| 2015 | Distributional determinants of learning argument structure constructions in first and second language
Yevgen Matusevych, Afra Alishahi, Ad Backus |
CogSci | 2 |
| 2014 | The impact of emerging knowledge of linguistic structure on word learning
Eva van den Bemd, Afra Alishahi, Maria Mos |
CogSci | 2 |
| 2014 | Isolating second language learning factors in a computational study of bilingual construction acquisition
Yevgen Matusevych, Afra Alishahi, Ad Backus |
CogSci | 2 |
| 2013 | Automatic generation of naturalistic child-adult interaction data
Yevgen Matusevych, Afra Alishahi, Paul Vogt |
CogSci | 2 |
| 2012 | Concurrent Acquisition of Word Meaning and Lexical Categories
Afra Alishahi, Grzegorz Chrupala |
EMNLP-CoNLL | 1 |
| 2011 | The Onset of Syntactic Bootstrapping in Word Learning: Evidence from a Computational Study
Afra Alishahi, Pirita Pyykkönen |
CogSci | 1 |
| 2010 | Online Entropy-Based Model of Lexical Category Acquisition
Grzegorz Chrupala, Afra Alishahi |
CoNLL | 2 |
| 2008 | Fast Mapping in Word Learning: What Probabilities Tell Us
Afra Alishahi, Afsaneh Fazly, Suzanne Stevenson |
CoNLL | 1 |