VLDB 2026 Research / reviewers in the wild / expert
Jasmijn Bastings
dblp:146/3824
· DBLP profile ↗
13ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0002-5445-4417ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Amplifying Trans and Nonbinary Voices: A Community-Centred Harm Taxonomy for LLMsabstractWe explore large language model (LLM) responses that may negatively impact the transgender and nonbinary (TGNB) community and introduce the Transing Transformers Toolkit, T^3, which provides resources for identifying such harmful response behaviors. The heart of T^3 is a community-centred taxonomy of harms, developed in collaboration with the TGNB community, which we complement with, amongst other guidance, suggested heuristics for evaluation. To develop the taxonomy, we adopted a multi-method approach that included surveys and focus groups with community experts. The contribution highlights the importance of community-centred approaches in mitigating harm, and outlines pathways for LLM developers to improve how their models handle TGNB-related topics. Eddie L. Ungless, Sunipa Dev, Cynthia L. Bennett, Rebecca Gulotta, Jasmijn Bastings, Remi Denton |
ACL (1) | 5 |
| 2024 | MiTTenS: A Dataset for Evaluating Gender MistranslationabstractTranslation systems, including foundation models capable of translation, can produce errors that result in gender mistranslations, and such errors create potential for harm.To measure the extent of such potential harms when translating into and out of English, we introduce a dataset, MiTTenS 1 , covering 26 languages from a variety of language families and scripts, including several traditionally underrepresented in digital resources.The dataset is constructed with handcrafted passages that target known failure patterns, longer synthetically generated passages, and natural passages sourced from multiple domains.We demonstrate the usefulness of the dataset by evaluating both neural machine translation systems and foundation models, and show that all systems exhibit gender mistranslation and potential harm, even in high resource languages. Kevin Robinson, Sneha Reddy Kudugunta, Romi Stella, Sunipa Dev, Jasmijn Bastings |
EMNLP | 5 |
| 2023 | Dissecting Recall of Factual Associations in Auto-Regressive Language ModelsabstractTransformer-based language models (LMs) are known to capture factual knowledge in their parameters.While previous work looked into where factual associations are stored, only little is known about how they are retrieved internally during inference.We investigate this question through the lens of information flow.Given a subject-relation query, we study how the model aggregates information about the subject and relation to predict the correct attribute.With interventions on attention edges, we first identify two critical points where information propagates to the prediction: one from the relation positions followed by another from the subject positions.Next, by analyzing the information at these points, we unveil a three-step internal mechanism for attribute extraction.First, the representation at the lastsubject position goes through an enrichment process, driven by the early MLP sublayers, to encode many subject-related attributes.Second, information from the relation propagates to the prediction.Third, the prediction representation "queries" the enriched subject to extract the attribute.Perhaps surprisingly, this extraction is typically done via attention heads, which often encode subject-attribute mappings in their parameters.Overall, our findings introduce a comprehensive view of how factual associations are stored and extracted internally in LMs, facilitating future research on knowledge localization and editing. 1 Mor Geva, Jasmijn Bastings, Katja Filippova, Amir Globerson |
EMNLP | 2 |
| 2023 | Scaling Vision Transformers to 22 Billion ParametersabstractThe scaling of Transformers has driven breakthrough capabilities for language models. At present, the largest large language models (LLMs) contain upwards of 100B parameters. Vision Transformers (ViT) have introduced the same architecture to image and video modelling, but these have not yet been successfully scaled to nearly the same degree; the largest dense ViT contains 4B parameters (Chen et al., 2022). We present a recipe for highly efficient and stable training of a 22B-parameter ViT (ViT-22B) and perform a wide variety of experiments on the resulting model. When evaluated on downstream tasks (often with a lightweight linear model on frozen features), ViT-22B demonstrates increasing performance with scale. We further observe other interesting benefits of scale, including an improved tradeoff between fairness and performance, state-of-the-art alignment to human visual perception in terms of shape/texture bias, and improved robustness. ViT-22B demonstrates the potential for "LLM-like" scaling in vision, and provides key steps towards getting there. Mostafa Dehghani 0001, Josip Djolonga, Basil Mustafa, Piotr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Steiner 0001, Mathilde Caron, Robert Geirhos, Ibrahim Alabdulmohsin, Rodolphe Jenatton, Lucas Beyer, Michael Tschannen, Anurag Arnab, Xiao Wang 0038, Carlos Riquelme, Matthias Minderer, Joan Puigcerver, Utku Evci, Sjoerd van Steenkiste, Gamaleldin F. Elsayed, Aravindh Mahendran, Fisher Yu 0001, Avital Oliver, Fantine Huot, Jasmijn Bastings, Mark Collier, Alexey A. Gritsenko, Vighnesh Birodkar, Cristina Nader Vasconcelos, Yi Tay, Thomas Mensink, Alexander Kolesnikov 0003, Filip Pavetic, Dustin Tran, Thomas Kipf, Mario Lucic, Xiaohua Zhai, Daniel Keysers, Jeremiah J. Harmsen, Neil Houlsby |
ICML | 27 |
| 2023 | Diagnosing AI Explanation Methods with Folk Concepts of Behavior
Alon Jacovi, Jasmijn Bastings, Sebastian Gehrmann, Yoav Goldberg, Katja Filippova |
J. Artif. Intell. Res. | 2 |
| 2023 | Scaling Up Models and Data with t5x and seqioabstractScaling up training datasets and model parameters have benefited neural network-based language models, but also present challenges like distributed compute, input data bottlenecks and reproducibility of results. We introduce two simple and scalable software libraries that simplify these issues: t5x enables training large language models at scale, while seqio enables reproducible input and evaluation pipelines. These open-source libraries have been used to train models with hundreds of billions of parameters on multi-terabyte datasets. Configurations and instructions for T5-like and GPT-like models are also provided. The libraries can be found at https://github.com/google-research/t5x and https://github.com/google/seqio. Adam Roberts, Hyung Won Chung, Anselm Levskaya, Daniel Andor, Sharan Narang, Brian Lester, Colin Gaffney, Afroz Mohiuddin, Curtis Hawthorne, Aitor Lewkowycz, Alex Salcianu, Marc van Zee, Jacob Austin, Sebastian Goodman, Livio B. Soares, Haitang Hu, Sasha Tsvyashchenko, Aakanksha Chowdhery, Jasmijn Bastings, Jannis Bulian, Xavier Garcia, Jianmo Ni, Kathleen Kenealy, Kehang Han, Michelle Casbon, Jonathan H. Clark, Stephan Lee, Dan Garrette, James Lee-Thorp, Colin Raffel, Noam Shazeer, Marvin Ritter, Maarten Bosma, Alexandre Tachard Passos, Jeremy Maitin-Shepard, Noah Fiedel, Mark Omernick, Brennan Saeta, Ryan Sepassi, Alexander Spiridonov, Joshua Newlan, Andrea Gesmundo |
J. Mach. Learn. Res. | 21 |
| 2022 | "Will You Find These Shortcuts?" A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text ClassificationabstractFeature attribution a.k.a.input salience methods which assign an importance score to a feature are abundant but may produce surprisingly different results for the same model on the same input.While differences are expected if disparate definitions of importance are assumed, most methods claim to provide faithful attributions and point at the features most relevant for a model's prediction.Existing work on faithfulness evaluation is not conclusive and does not provide a clear answer as to how different methods are to be compared.Focusing on text classification and the model debugging scenario, our main contribution is a protocol for faithfulness evaluation that makes use of partially synthetic data to obtain ground truth for feature importance ranking.Following the protocol, we do an in-depth analysis of four standard salience method classes on a range of datasets and lexical shortcuts for BERT and LSTM models.We demonstrate that some of the most popular method configurations provide poor results even for simple shortcuts while a method judged to be too simplistic works remarkably well for BERT. Jasmijn Bastings, Sebastian Ebert, Polina Zablotskaia, Anders Sandholm 0001, Katja Filippova |
EMNLP | 1 |
| 2022 | Autoregressive Diffusion Models
Emiel Hoogeboom, Alexey A. Gritsenko, Jasmijn Bastings, Ben Poole, Rianne van den Berg, Tim Salimans |
ICLR | 3 |
| 2022 | The MultiBERTs: BERT Reproductions for Robustness Analysis
Thibault Sellam, Steve Yadlowsky, Ian Tenney, Jason Wei, Naomi Saphra, Alexander D'Amour, Tal Linzen, Jasmijn Bastings, Iulia Turc, Jacob Eisenstein, Dipanjan Das 0001, Ellie Pavlick |
ICLR | 8 |
| 2021 | We Need To Talk About Random SplitsabstractGorman and Bedrick (2019) argued for using random splits rather than standard splits in NLP experiments.We argue that random splits, like standard splits, lead to overly optimistic performance estimates.We can also split data in biased or adversarial ways, e.g., training on short sentences and evaluating on long ones.Biased sampling has been used in domain adaptation to simulate real-world drift; this is known as the covariate shift assumption.In NLP, however, even worst-case splits, maximizing bias, often under-estimate the error observed on new samples of in-domain data, i.e., the data that models should minimally generalize to at test time.This invalidates the covariate shift assumption.Instead of using multiple random splits, future benchmarks should ideally include multiple, independent test sets instead; if infeasible, we argue that multiple biased splits leads to more realistic performance estimates than multiple random splits. Anders Søgaard, Sebastian Ebert, Jasmijn Bastings, Katja Filippova |
EACL | 3 |
| 2019 | Interpretable Neural Predictions with Differentiable Binary VariablesabstractThe success of neural networks comes hand in hand with a desire for more interpretability.We focus on text classifiers and make them more interpretable by having them provide a justification-a rationale-for their predictions.We approach this problem by jointly training two neural network models: a latent model that selects a rationale (i.e. a short and informative part of the input text), and a classifier that learns from the words in the rationale alone.Previous work proposed to assign binary latent masks to input positions and to promote short selections via sparsityinducing penalties such as L 0 regularisation.We propose a latent model that mixes discrete and continuous behaviour allowing at the same time for binary selections and gradient-based training without REINFORCE.In our formulation, we can tractably compute the expected value of penalties such as L 0 , which allows us to directly optimise the model towards a prespecified text selection rate.We show that our approach is competitive with previous work on rationale extraction, and explore further uses in attention mechanisms. Jasmijn Bastings, Wilker Aziz, Ivan Titov 0001 |
ACL (1) | 1 |
| 2017 | Graph Convolutional Encoders for Syntax-aware Neural Machine TranslationabstractWe present a simple and effective approach to incorporating syntactic structure into neural attention-based encoderdecoder models for machine translation.We rely on graph-convolutional networks (GCNs), a recent class of neural networks developed for modeling graph-structured data.Our GCNs use predicted syntactic dependency trees of source sentences to produce representations of words (i.e.hidden states of the encoder) that are sensitive to their syntactic neighborhoods.GCNs take word representations as input and produce word representations as output, so they can easily be incorporated as layers into standard encoders (e.g., on top of bidirectional RNNs or convolutional neural networks).We evaluate their effectiveness with English-German and English-Czech translation experiments for different types of encoders and observe substantial improvements over their syntax-agnostic versions in all the considered setups. Jasmijn Bastings, Ivan Titov 0001, Wilker Aziz, Diego Marcheggiani, Khalil Sima'an |
EMNLP | 1 |
| 2014 | All Fragments Count in Parser Evaluation
Jasmijn Bastings, Khalil Sima'an |
LREC | 1 |