Grzegorz Chrupala

dblp:19/1379 · DBLP profile ↗
← Back
36ranked-venue papers
11as first author
12since 2021 · last 2026
0000-0001-9498-6912ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 36 · 11 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Mechanistic Interpretability Meets Cognitive Linguistics: Modelling Locative Image Schemas in the Circuit Framework
Mattia Proietti, Afra Alishahi, Grzegorz Chrupala, Alessandro Lenci
LREC3
2025 On the reliability of feature attribution methods for speech classification
abstract
As the capabilities of large-scale pre-trained models evolve, understanding the determinants of their outputs becomes more important. Feature attribution aims to reveal which parts of the input elements contribute the most to model outputs. In speech processing, the unique characteristics of the input signal make the application of feature attribution methods challenging. We study how factors such as input type and aggregation and perturbation timespan impact the reliability of standard feature attribution methods, and how these factors interact with characteristics of each classification task. We find that standard approaches to feature attribution are generally unreliable when applied to the speech domain, with the exception of word-aligned perturbation methods when applied to word-based classification tasks.
Gaofei Shen, Hosein Mohebbi, Arianna Bisazza, Afra Alishahi, Grzegorz Chrupala
INTERSPEECH5
2025 QE4PE: Word-level Quality Estimation for Human Post-Editing
Gabriele Sarti, Vilém Zouhar, Grzegorz Chrupala, Ana Guerberof Arenas, Malvina Nissim, Arianna Bisazza
Trans. Assoc. Comput. Linguistics3
2024 Quantifying the Plausibility of Context Reliance in Neural Machine Translation
abstract
Establishing whether language models can use contextual information in a human-plausible way is important to ensure their safe adoption in real-world settings. However, the questions of $\textit{when}$ and $\textit{which parts}$ of the context affect model generations are typically tackled separately, and current plausibility evaluations are practically limited to a handful of artificial benchmarks. To address this, we introduce $\textbf{P}$lausibility $\textbf{E}$valuation of $\textbf{Co}$ntext $\textbf{Re}$liance (PECoRe), an end-to-end interpretability framework designed to quantify context usage in language models' generations. Our approach leverages model internals to (i) contrastively identify context-sensitive target tokens in generated texts and (ii) link them to contextual cues justifying their prediction. We use PECoRe to quantify the plausibility of context-aware machine translation models, comparing model rationales with human annotations across several discourse-level phenomena. Finally, we apply our method to unannotated model translations to identify context-mediated predictions and highlight instances of (im)plausible context usage throughout generation.
Gabriele Sarti, Grzegorz Chrupala, Malvina Nissim, Arianna Bisazza
ICLR2
2024 Encoding of lexical tone in self-supervised models of spoken language
abstract
Gaofei Shen, Michaela Watkins, Afra Alishahi, Arianna Bisazza, Grzegorz Chrupała. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Gaofei Shen, Michaela Watkins, Afra Alishahi, Arianna Bisazza, Grzegorz Chrupala
NAACL-HLT5
2023 Quantifying Context Mixing in Transformers
abstract
Self-attention weights and their transformed variants have been the main source of information for analyzing token-to-token interactions in Transformer-based models. But despite their ease of interpretation, these weights are not faithful to the models’ decisions as they are only one part of an encoder, and other components in the encoder layer can have considerable impact on information mixing in the output representations. In this work, by expanding the scope of analysis to the whole encoder block, we propose Value Zeroing, a novel context mixing score customized for Transformers that provides us with a deeper understanding of how information is mixed at each encoder layer. We demonstrate the superiority of our context mixing score over other analysis methods through a series of complementary evaluations with different viewpoints based on linguistically informed rationales, probing, and faithfulness analysis.
Hosein Mohebbi, Willem H. Zuidema, Grzegorz Chrupala, Afra Alishahi
EACL3
2023 Homophone Disambiguation Reveals Patterns of Context Mixing in Speech Transformers
abstract
Transformers have become a key architecture in speech processing, but our understanding of how they build up representations of acoustic and linguistic structure is limited.In this study, we address this gap by investigating how measures of 'context-mixing' developed for text models can be adapted and applied to models of spoken language.We identify a linguistic phenomenon that is ideal for such a case study: homophony in French (e.g.livre vs livres), where a speech recognition model has to attend to syntactic cues such as determiners and pronouns in order to disambiguate spoken words with identical pronunciations and transcribe them while respecting grammatical agreement.We perform a series of controlled experiments and probing analyses on Transformer-based speech models.Our findings reveal that representations in encoder-only models effectively incorporate these cues to identify the correct transcription, whereas encoders in encoder-decoder models mainly relegate the task of capturing contextual dependencies to decoder modules. 1
Hosein Mohebbi, Grzegorz Chrupala, Willem H. Zuidema, Afra Alishahi
EMNLP2
2023 Wave to Syntax: Probing spoken language models for syntax
abstract
Understanding which information is encoded in deep models of spoken and written language has been the focus of much research in recent years, as it is crucial for debugging and improving these architectures. Most previous work has focused on probing for speaker characteristics, acoustic and phonological information in models of spoken language, and for syntactic information in models of written language. Here we focus on the encoding of syntax in several self-supervised and visually grounded models of spoken language. We employ two complementary probing methods, combined with baselines and reference representations to quantify the degree to which syntactic structure is encoded in the activations of the target models. We show that syntax is captured most prominently in the middle layers of the networks, and more explicitly within models with more parameters.
Gaofei Shen, Afra Alishahi, Arianna Bisazza, Grzegorz Chrupala
INTERSPEECH4
2022 Cyberbullying Classifiers are Sensitive to Model-Agnostic Perturbations
abstract
A limited amount of studies investigates the role of model-agnostic adversarial behavior in toxic content classification. As toxicity classifiers predominantly rely on lexical cues, (deliberately) creative and evolving language-use can be detrimental to the utility of current corpora and state-of-the-art models when they are deployed for content moderation. The less training data is available, the more vulnerable models might become. This study is, to our knowledge, the first to investigate the effect of adversarial behavior and augmentation for cyberbullying detection. We demonstrate that model-agnostic lexical substitutions significantly hurt classifier performance. Moreover, when these perturbed samples are used for augmentation, we show models become robust against word-level perturbations at a slight trade-off in overall task performance. Augmentations proposed in prior work on toxicity prove to be less effective. Our results underline the need for such evaluations in online harm areas with small corpora.
Chris Emmery, Ákos Kádár, Grzegorz Chrupala, Walter Daelemans
LREC3
2022 Visually Grounded Models of Spoken Language: A Survey of Datasets, Architectures and Evaluation Techniques
abstract
This survey provides an overview of the evolution of visually grounded models of spoken language over the last 20 years. Such models are inspired by the observation that when children pick up a language, they rely on a wide range of indirect and noisy clues, crucially including signals from the visual modality co-occurring with spoken utterances. Several fields have made important contributions to this approach to modeling or mimicking the process of learning language: Machine Learning, Natural Language and Speech Processing, Computer Vision and Cognitive Science. The current paper brings together these contributions in order to provide a useful introduction and overview for practitioners in all these areas. We discuss the central research questions addressed, the timeline of developments, and the datasets which enabled much of this work. We then summarize the main modeling architectures and offer an exhaustive overview of the evaluation metrics and analysis techniques.
Grzegorz Chrupala
J. Artif. Intell. Res.1
2022 Learning English with Peppa Pig
abstract
Abstract Recent computational models of the acquisition of spoken language via grounding in perception exploit associations between spoken and visual modalities and learn to represent speech and visual data in a joint vector space. A major unresolved issue from the point of ecological validity is the training data, typically consisting of images or videos paired with spoken descriptions of what is depicted. Such a setup guarantees an unrealistically strong correlation between speech and the visual data. In the real world the coupling between the linguistic and the visual modality is loose, and often confounded by correlations with non-semantic aspects of the speech signal. Here we address this shortcoming by using a dataset based on the children’s cartoon Peppa Pig. We train a simple bi-modal architecture on the portion of the data consisting of dialog between characters, and evaluate on segments containing descriptive narrations. Despite the weak and confounded signal in this training data, our model succeeds at learning aspects of the visual semantics of spoken language.
Mitja Nikolaus, Afra Alishahi, Grzegorz Chrupala
Trans. Assoc. Comput. Linguistics3
2021 Adversarial Stylometry in the Wild: Transferable Lexical Substitution Attacks on Author Profiling
abstract
Written language contains stylistic cues that can be exploited to automatically infer a variety of potentially sensitive author information.Adversarial stylometry intends to attack such models by rewriting an author's text.Our research proposes several components to facilitate deployment of these adversarial attacks in the wild, where neither data nor target models are accessible.We introduce a transformerbased extension of a lexical replacement attack, and show it achieves high transferability when trained on a weakly labeled corpusdecreasing target model performance below chance.While not completely inconspicuous, our more successful attacks also prove notably less detectable by humans.Our framework therefore provides a promising direction for future privacy-preserving adversarial attacks.
Chris Emmery, Ákos Kádár, Grzegorz Chrupala
EACL3
2020 Analyzing analytical methods: The case of phonology in neural models of spoken language
abstract
Given the fast development of analysis techniques for NLP and speech processing systems, few systematic studies have been conducted to compare the strengths and weaknesses of each method. As a step in this direction we study the case of representations of phonology in neural network models of spoken language. We use two commonly applied analytical techniques, diagnostic classifiers and representational similarity analysis, to quantify to what extent neural activation patterns encode phonemes and phoneme sequences. We manipulate two factors that can affect the outcome of analysis. First, we investigate the role of learning by comparing neural activations extracted from trained versus randomly-initialized models. Second, we examine the temporal scope of the activations by probing both local activations corresponding to a few milliseconds of the speech signal, and global activations pooled over the whole utterance. We conclude that reporting analysis results with randomly initialized models is crucial, and that global-scope methods tend to yield more consistent results and we recommend their use as a complement to local-scope diagnostic methods.
Grzegorz Chrupala, Bertrand Higy, Afra Alishahi
ACL1
2020 Learning to Understand Child-directed and Adult-directed Speech
abstract
Speech directed to children differs from adult-directed speech in linguistic aspects such as repetition, word choice, and sentence length, as well as in aspects of the speech signal itself, such as prosodic and phonemic variation. Human language acquisition research indicates that child-directed speech helps language learners. This study explores the effect of child-directed speech when learning to extract semantic information from speech directly. We compare the task performance of models trained on adult-directed speech (ADS) and child-directed speech (CDS). We find indications that CDS helps in the initial stages of learning, but eventually, models trained on ADS reach comparable task performance, and generalize better. The results suggest that this is at least partially due to linguistic rather than acoustic properties of the two registers, as we see the same pattern when looking at models trained on acoustically comparable synthetic speech.
Lieke Gelderloos, Grzegorz Chrupala, Afra Alishahi
ACL2
2019 Symbolic Inductive Bias for Visually Grounded Learning of Spoken Language
abstract
A widespread approach to processing spoken language is to first automatically transcribe it into text. An alternative is to use an end-to-end approach: recent works have proposed to learn semantic embeddings of spoken language from images with spoken captions, without an intermediate transcription step. We propose to use multitask learning to exploit existing transcribed speech within the end-to-end setting. We describe a three-task architecture which combines the objectives of matching spoken captions with corresponding images, speech with text, and text with images. We show that the addition of the speech/text task leads to substantial performance improvements on image retrieval when compared to training the speech/image task in isolation. We conjecture that this is due to a strong inductive bias transcribed speech provides to the model, and offer supporting evidence for this.
Grzegorz Chrupala
ACL (1)1
2019 Correlating Neural and Symbolic Representations of Language
abstract
Analysis methods which enable us to better understand the representations and functioning of neural models of language are increasingly needed as deep learning becomes the dominant approach in NLP.Here we present two methods based on Representational Similarity Analysis (RSA) and Tree Kernels (TK) which allow us to directly quantify how strongly the information encoded in neural activation patterns corresponds to information represented by symbolic structures such as syntax trees.We first validate our methods on the case of a simple synthetic language for arithmetic expressions with clearly defined syntax and semantics, and show that they exhibit the expected pattern of results.We then apply our methods to correlate neural representations of English sentences with their constituency parse trees.
Grzegorz Chrupala, Afra Alishahi
ACL (1)1
2019 Analyzing and interpreting neural networks for NLP: A report on the first BlackboxNLP workshop
abstract
Abstract The Empirical Methods in Natural Language Processing (EMNLP) 2018 workshop BlackboxNLP was dedicated to resources and techniques specifically developed for analyzing and understanding the inner-workings and representations acquired by neural models of language. Approaches included: systematic manipulation of input to neural networks and investigating the impact on their performance, testing whether interpretable knowledge can be decoded from intermediate representations acquired by neural networks, proposing modifications to neural network architectures to make their knowledge state or generated output more explainable, and examining the performance of networks on simplified or formal languages. Here we review a number of representative studies in each category.
Afra Alishahi, Grzegorz Chrupala, Tal Linzen
Nat. Lang. Eng.2
2018 Style Obfuscation by Invariance
abstract
The task of obfuscating writing style using sequence models has previously been investigated under the framework of obfuscation-by-transfer, where the input text is explicitly rewritten in another style. A side effect of this framework are the frequent major alterations to the semantic content of the input. In this work, we propose obfuscation-by-invariance, and investigate to what extent models trained to be explicitly style-invariant preserve semantics. We evaluate our architectures in parallel and non-parallel settings, and compare automatic and human evaluations on the obfuscated sentences. Our experiments show that the performance of a style classifier can be reduced to chance level, while the output is evaluated to be of equal quality to models applying style-transfer. Additionally, human evaluation indicates a trade-off between the level of obfuscation and the observed quality of the output in terms of meaning preservation and grammaticality.
Chris Emmery, Enrique Manjavacas, Grzegorz Chrupala
COLING3
2018 Revisiting the Hierarchical Multiscale LSTM
abstract
Hierarchical Multiscale LSTM (Chung et. al., 2016) is a state-of-the-art language model that learns interpretable structure from character-level input. Such models can provide fertile ground for (cognitive) computational linguistics studies. However, the high complexity of the architecture, training and implementations might hinder its applicability. We provide a detailed reproduction and ablation study of the architecture, shedding light on some of the potential caveats of re-purposing complex deep-learning architectures. We further show that simplifying certain aspects of the architecture can in fact improve its performance. We also investigate the linguistic units (segments) learned by various levels of the model, and argue that their quality does not correlate with the overall performance of the model on language modeling.
Ákos Kádár, Marc-Alexandre Côté, Grzegorz Chrupala, Afra Alishahi
COLING3
2018 Lessons Learned in Multilingual Grounded Language Learning
abstract
Recent work has shown how to learn better visual-semantic embeddings by leveraging image descriptions in more than one language.Here, we investigate in detail which conditions affect the performance of this type of grounded language learning model.We show that multilingual training improves over bilingual training, and that low-resource languages benefit from training with higher-resource languages.We demonstrate that a multilingual model can be trained equally well on either translations or comparable sentence pairs, and that annotating the same set of images in multiple language enables further improvements via an additional caption-caption ranking objective.
Ákos Kádár, Desmond Elliott, Marc-Alexandre Côté, Grzegorz Chrupala, Afra Alishahi
CoNLL4
2017 Representations of language in a model of visually grounded speech signal
abstract
We present a visually grounded model of speech perception which projects spoken utterances and images to a joint semantic space.We use a multi-layer recurrent highway network to model the temporal nature of spoken speech, and show that it learns to extract both form and meaningbased linguistic knowledge from the input signal.We carry out an in-depth analysis of the representations used by different components of the trained model and show that encoding of semantic aspects tends to become richer as we go up the hierarchy of layers, whereas encoding of formrelated aspects of the language input tends to initially increase and then plateau or decrease.
Grzegorz Chrupala, Lieke Gelderloos, Afra Alishahi
ACL (1)1
2017 Encoding of phonology in a recurrent neural model of grounded speech
abstract
We study the representation and encoding of phonemes in a recurrent neural network model of grounded speech.We use a model which processes images and their spoken descriptions, and projects the visual and auditory representations into the same semantic space.We perform a number of analyses on how information about individual phonemes is encoded in the MFCC features extracted from the speech signal, and the activations of the layers of the model.Via experiments with phoneme decoding and phoneme discrimination we show that phoneme representations are most salient in the lower layers of the model, where low-level signals are processed at a fine-grained level, although a large amount of phonological information is retain at the top recurrent layer.We further find out that the attention mechanism following the top recurrent layer significantly attenuates encoding of phonology and makes the utterance embeddings much more invariant to synonymy.Moreover, a hierarchical clustering of phoneme representations learned by the network shows an organizational structure of phonemes similar to those proposed in linguistics.
Afra Alishahi, Marie Barking, Grzegorz Chrupala
CoNLL3
2017 Representation of Linguistic Form and Function in Recurrent Neural Networks
abstract
We present novel methods for analyzing the activation patterns of recurrent neural networks from a linguistic point of view and explore the types of linguistic structure they learn. As a case study, we use a standard standalone language model, and a multi-task gated recurrent network architecture consisting of two parallel pathways with shared word embeddings: The Visual pathway is trained on predicting the representations of the visual scene corresponding to an input sentence, and the Textual pathway is trained to predict the next word in the same sentence. We propose a method for estimating the amount of contribution of individual tokens in the input to the final prediction of the networks. Using this method, we show that the Visual pathway pays selective attention to lexical categories and grammatical functions that carry semantic information, and learns to treat word types differently depending on their grammatical function and their position in the sequential structure of the sentence. In contrast, the language models are comparatively more sensitive to words with a syntactic function. Further analysis of the most informative n-gram contexts for each model shows that in comparison with the Visual pathway, the language models react more strongly to abstract contexts that represent syntactic constructions.
Ákos Kádár, Grzegorz Chrupala, Afra Alishahi
Comput. Linguistics2
2016 From phonemes to images: levels of representation in a recurrent neural model of visually-grounded language learning
abstract
We present a model of visually-grounded language learning based on stacked gated recurrent neural networks which learns to predict visual features given an image description in the form of a sequence of phonemes. The learning task resembles that faced by human language learners who need to discover both structure and meaning from noisy and ambiguous data across modalities. We show that our model indeed learns to predict features of the visual context given phonetically transcribed image descriptions, and show that it represents linguistic information in a hierarchy of levels: lower layers in the stack are comparatively more sensitive to form, whereas higher layers are more sensitive to meaning.
Lieke Gelderloos, Grzegorz Chrupala
COLING2
2016 Multimodal Semantic Learning from Child-Directed Input
abstract
Angeliki Lazaridou, Grzegorz Chrupała, Raquel Fernández, Marco Baroni. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Angeliki Lazaridou, Grzegorz Chrupala, Raquel Fernández, Marco Baroni
HLT-NAACL2
2014 RelationFactory: A Fast, Modular and Effective System for Knowledge Base Population
abstract
Benjamin Roth, Tassilo Barth, Grzegorz Chrupała, Martin Gropp, Dietrich Klakow. Proceedings of the Demonstrations at the 14th Conference of the European Chapter of the Association for Computational Linguistics. 2014.
Benjamin Roth 0001, Tassilo Barth, Grzegorz Chrupala, Martin Gropp, Dietrich Klakow
EACL3
2014 Semantic approaches to software component retrieval with English queries
Huijing Deng, Grzegorz Chrupala
LREC2
2013 Elephant: Sequence Labeling for Word and Sentence Segmentation
abstract
Tokenization is widely regarded as a solved problem due to the high accuracy that rulebased tokenizers achieve.But rule-based tokenizers are hard to maintain and their rules language specific.We show that highaccuracy word and sentence segmentation can be achieved by using supervised sequence labeling on the character level combined with unsupervised feature learning.We evaluated our method on three languages and obtained error rates of 0.27 ‰ (English), 0.35 ‰ (Dutch) and 0.76 ‰ (Italian) for our best models.
Kilian Evang, Valerio Basile, Grzegorz Chrupala, Johan Bos
EMNLP3
2012 Learning from evolving data streams: online triage of bug reports
Grzegorz Chrupala
EACL1
2012 Concurrent Acquisition of Word Meaning and Lexical Categories
Afra Alishahi, Grzegorz Chrupala
EMNLP-CoNLL2
2011 Efficient induction of probabilistic word classes with LDA
Grzegorz Chrupala
IJCNLP1
2010 Online Entropy-Based Model of Lexical Category Acquisition
Grzegorz Chrupala, Afra Alishahi
CoNLL1
2010 A Named Entity Labeler for German: Exploiting Wikipedia and Distributional Clusters
Grzegorz Chrupala, Dietrich Klakow
LREC1
2008 Learning Morphology with Morfette
Grzegorz Chrupala, Georgiana Dinu, Josef van Genabith
LREC1
2006 Using Machine-Learning to Assign Function Labels to Parser Output for Spanish
Grzegorz Chrupala, Josef van Genabith
ACL1
2004 Hierarchical Recognition of Propositional Arguments with Perceptrons
Xavier Carreras, Lluís Màrquez, Grzegorz Chrupala
CoNLL3