VLDB 2026 Research / reviewers in the wild / expert
Mariya Toneva
dblp:160/4677
· DBLP profile ↗
19ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0002-2407-9871ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 3 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | What Brain Data Adds to Language Model TrainingabstractBrain-tuning language models (LMs)—fine-tuning LMs to predict brain recordings elicited by linguistic stimuli—has been proposed as a promising way to align LMs closer to the human brain, with recent work reporting gains on a small number of downstream tasks. However, it remains unclear what benefits brain data provide beyond those obtainable from further training on the same underlying linguistic input, and whether such benefits generalize across tasks. Here, we present a comprehensive evaluation of jointly-tuned LMs, trained on both brain recordings and text-based stimuli, brain-tuned LMs and LMs tuned only on text-based stimuli (i.e., stimulus-tuned LMs). We compare models across a diverse suite of downstream linguistic tasks. We find that jointly-tuned LMs outperform other fine-tuned and pretrained models, and that brain-tuned LMs outperform stimulus-tuned LMs, demonstrating the richness of brain data as an additional training signal for LMs. Gabriele Merlin, Omer Moussa, Mariya Toneva |
CoNLL | 3 |
| 2025 | Improving Semantic Understanding in Speech Language Models via Brain-tuningabstractSpeech language models align with human brain responses to natural language to an impressive degree. However, current models rely heavily on low-level speech features, indicating they lack brain-relevant semantics which limits their utility as model organisms of semantic processing in the brain. In this work, we address this limitation by inducing brain-relevant bias directly into the models via fine-tuning with fMRI recordings of people listening to natural stories--a process we name brain-tuning. After testing it on 3 different pretrained model families, we show that brain-tuning not only improves overall alignment with new brain recordings in semantic language regions, but also reduces the reliance on low-level speech features for this alignment. Excitingly, we further show that brain-tuning leads to 1) consistent improvements in performance on semantic downstream tasks and 2) a representational space with increased semantic preference. Our results provide converging evidence, for the first time, that incorporating brain signals into the training of language models improves the models’ semantic understanding. We make the code available at https://github.com/bridge-ai-neuro/brain-tuning. Omer Moussa, Dietrich Klakow, Mariya Toneva |
ICLR | 3 |
| 2025 | Hints Help Finding and Fixing Bugs Differently in Python and Text-Based Program RepresentationsabstractWith the recent advances in AI programming assistants such as GitHub Copilot, programming is not limited to classical programming languages anymore-programming tasks can also be expressed and solved by end-users in natural text. Despite the availability of this new programming modality, users still face difficulties with algorithmic understanding and program debugging. One promising approach to support end-users is to provide hints to help them find and fix bugs while forming and improving their programming capabilities. While it is plausible that hints can help, it is unclear which type of hint is helpful and how this depends on program representations (classic source code or a textual representation) and the user's capability of understanding the algorithmic task. To understand the role of hints in this space, we conduct a large-scale crowd-sourced study involving 753 participants investigating the effect of three types of hints (test cases, conceptual, and detailed), across two program representations (Python and text-based), and two groups of users (with clear understanding or confusion about the algorithmic task). We find that the program representation (Python vs. text) has a significant influence on the users' accuracy at finding and fixing bugs. Surprisingly, users are more accurate at finding and fixing bugs when they see the program in natural text. Hints are generally helpful in improving accuracy, but different hints help differently depending on the program representation and the user's understanding of the algorithmic task. These findings have implications for designing next-generation programming tools that provide personalized support to users, for example, by adapting the programming modality and providing hints with respect to the user's skill level and understanding. Ruchit Rawal, Victor-Alexandru Padurean, Sven Apel, Adish Singla, Mariya Toneva |
ICSE | 5 |
| 2025 | Brain-tuned Speech Models Better Reflect Speech Processing Stages in the BrainabstractPretrained self-supervised speech models excel in speech tasks but do not reflect the hierarchy of human speech processing, as they encode rich semantics in middle layers and poor semantics in late layers. Recent work showed that brain-tuning (fine-tuning models using human brain recordings) improves speech models' semantic understanding. Here, we examine how well brain-tuned models further reflect the brain's intermediate stages of speech processing. We find that late layers of brain-tuned models substantially improve over pretrained models in their alignment with semantic language regions. Further layer-wise probing reveals that early layers remain dedicated to low-level acoustic features, while late layers become the best at complex high-level tasks. These findings show that brain-tuned models not only perform better but also exhibit a well-defined hierarchical processing going from acoustic to semantic representations, making them better model organisms for human speech processing. Omer Moussa, Mariya Toneva |
INTERSPEECH | 2 |
| 2025 | Large Language Models as Model Organisms for Human Associative LearningabstractAssociative learning--forming links between co-occurring items--is fundamental to human cognition, reshaping internal representations in complex ways. Testing hypotheses on how representational changes occur in biological systems is challenging, but large language models (LLMs) offer a scalable alternative. Building on LLMs' in-context learning, we adapt a cognitive neuroscience associative learning paradigm and investigate how representations evolve across six models. Our initial findings reveal a non-monotonic pattern consistent with the Non-Monotonic Plasticity Hypothesis, with moderately similar items differentiating after learning. Leveraging the controllability of LLMs, we further show that this differentiation is modulated by the overlap of associated items with the broader vocabulary--a factor we term vocabulary interference, capturing how new associations compete with prior knowledge. We find that higher vocabulary interference amplifies differentiation, suggesting that representational change is influenced by both item similarity and global competition. Our findings position LLMs not only as powerful tools for studying representational dynamics in human-like learning systems, but also as accessible and general computational models for generating new hypotheses about the principles underlying memory reorganization in the brain. Camila Kolling, Vy A. Vo, Mariya Toneva |
NeurIPS | 3 |
| 2025 | Brain-tuning Improves Generalizability and Efficiency of Brain Alignment in Speech ModelsabstractPretrained language models are remarkably effective in aligning with human brain responses elicited by natural language stimuli, positioning them as promising model organisms for studying language processing in the brain. However, existing approaches for both estimating and improving this brain alignment are participant-dependent and highly affected by the amount of data available per participant, hindering both generalization to new participants and population-level analyses. In this work, we address these limitations by introducing a scalable, generalizable brain-tuning method, in which we fine-tune pretrained speech language models to jointly predict fMRI responses from multiple participants. We demonstrate that the resulting brain-tuned models exhibit strong individual brain alignment while generalizing across participants. Specifically, our method leads to 1) a 5-fold decrease in the amount of fMRI data needed to predict brain data from new participants, 2) up to a 50\% increase in the overall brain alignment, and 3) strong generalization to new unseen datasets. Furthermore, this multi-participant brain-tuning additionally improves downstream performance on semantic tasks, suggesting that training using brain data from multiple participants leads to more generalizable semantic representations. Taken together, these findings demonstrate a bidirectional benefit between neuroscience and AI, helping bridge the gap between the two fields. We make our code and models publicly available at https://github.com/bridge-ai-neuro/multi-brain-tuning. Omer Moussa, Mariya Toneva |
NeurIPS | 2 |
| 2024 | Speech language models lack important brain-relevant semanticsabstractDespite known differences between reading and listening in the brain, recent work has shown that text-based language models predict both text-evoked and speech-evoked brain activity to an impressive degree.This poses the question of what types of information language models truly predict in the brain.We investigate this question via a direct approach, in which we systematically remove specific lowlevel stimulus features (textual, speech, and visual) from language model representations to assess their impact on alignment with fMRI brain recordings during reading and listening.Comparing these findings with speech-based language models reveals starkly different effects of low-level features on brain alignment.While text-based models show reduced alignment in early sensory regions post-removal, they retain significant predictive power in late language regions.In contrast, speech-based models maintain strong alignment in early auditory regions even after feature removal but lose all predictive power in late language regions.These results suggest that speech-based models provide insights into additional information processed by early auditory regions, but caution is needed when using them to model processing in late language regions.We make our code publicly available.1 Subba Reddy Oota, Emin Çelik, Fatma Deniz, Mariya Toneva |
ACL (1) | 4 |
| 2024 | Is Deep Learning the Answer for Understanding Human Cognitive Dynamics?
John P. Spencer, Brenden M. Lake, Raul Grieben, Gregor Schöner, Mariya Toneva, Gina R. Kuperberg |
CogSci | 5 |
| 2024 | Language models and brains align due to more than next-word prediction and word-level informationabstractPretrained language models have been shown to significantly predict brain recordings of people comprehending language. Recent work suggests that the prediction of the next word is a key mechanism that contributes to this alignment. What is not yet understood is whether prediction of the next word is necessary for this observed alignment or simply sufficient, and whether there are other shared mechanisms or information that are similarly important. In this work, we take a step towards understanding the reasons for brain alignment via two simple perturbations in popular pretrained language models. These perturbations help us design contrasts that can control for different types of information. By contrasting the brain alignment of these differently perturbed models, we show that improvements in alignment with brain recordings are due to more than improvements in next-word prediction and word-level information. Gabriele Merlin, Mariya Toneva |
EMNLP | 2 |
| 2023 | Training language models to summarize narratives improves brain alignment
Khai Loong Aw, Mariya Toneva |
ICLR | 2 |
| 2023 | Joint processing of linguistic properties in brains and language modelsabstractLanguage models have been shown to be very effective in predicting brain recordings of subjects experiencing complex language stimuli. For a deeper understanding of this alignment, it is important to understand the correspondence between the detailed processing of linguistic information by the human brain versus language models. We investigate this correspondence via a direct approach, in which we eliminate information related to specific linguistic properties in the language model representations and observe how this intervention affects the alignment with fMRI brain recordings obtained while participants listened to a story. We investigate a range of linguistic properties (surface, syntactic, and semantic) and find that the elimination of each one results in a significant decrease in brain alignment. Specifically, we find that syntactic properties (i.e. Top Constituents and Tree Depth) have the largest effect on the trend of brain alignment across model layers. These findings provide clear evidence for the role of specific linguistic information in the alignment between brain and language models, and open new avenues for mapping the joint information processing in both systems. We make the code publicly available https://github.com/subbareddy248/lingprop-brain-alignment. Subba Reddy Oota, Manish Gupta 0001, Mariya Toneva |
NeurIPS | 3 |
| 2022 | Deep Learning for Brain Encoding and Decoding
Subba Reddy Oota, Jashn Arora, Manish Gupta 0001, Raju S. Bapi, Mariya Toneva |
CogSci | 5 |
| 2020 | Modeling Task Effects on Meaning Representation in the Brain via Zero-Shot MEG PredictionabstractHow meaning is represented in the brain is still one of the big open questions in neuroscience. Does a word (e.g., bird) always have the same representation, or does the task under which the word is processed alter its representation (answering can you eat it?" versuscan it fly?")? The brain activity of subjects who read the same word while performing different semantic tasks has been shown to differ across tasks. However, it is still not understood how the task itself contributes to this difference. In the current work, we study Magnetoencephalography (MEG) brain recordings of participants tasked with answering questions about concrete nouns. We investigate the effect of the task (i.e. the question being asked) on the processing of the concrete noun by predicting the millisecond-resolution MEG recordings as a function of both the semantics of the noun and the task. Using this approach, we test several hypotheses about the task-stimulus interactions by comparing the zero-shot predictions made by these hypotheses for novel tasks and nouns not seen during training. We find that incorporating the task semantics significantly improves the prediction of MEG recordings, across participants. The improvement occurs 475-550ms after the participants first see the word, which corresponds to what is considered to be the ending time of semantic processing for a word. These results suggest that only the end of semantic processing of a word is task-dependent, and pose a challenge for future research to formulate new hypotheses for earlier task effects as a function of the task and stimuli. Mariya Toneva, Otilia Stretcu, Barnabás Póczos, Leila Wehbe, Tom M. Mitchell |
NeurIPS | 1 |
| 2019 | An Empirical Study of Example Forgetting during Deep Neural Network Learning
Mariya Toneva, Alessandro Sordoni, Remi Tachet des Combes, Adam Trischler, Yoshua Bengio, Geoffrey J. Gordon |
ICLR (Poster) | 1 |
| 2019 | Inducing brain-relevant bias in natural language processing modelsabstractProgress in natural language processing (NLP) models that estimate representations of word sequences has recently been leveraged to improve the understanding of language processing in the brain. However, these models have not been specifically designed to capture the way the brain represents language meaning. We hypothesize that fine-tuning these models to predict recordings of brain activity of people reading text will lead to representations that encode more brain-activity-relevant language information. We demonstrate that a version of BERT, a recently introduced and powerful language model, can improve the prediction of brain activity after fine-tuning. We show that the relationship between language and brain activity learned by BERT during this fine-tuning transfers across multiple participants. We also show that, for some participants, the fine-tuned representations learned from both magnetoencephalography (MEG) and functional magnetic resonance imaging (fMRI) are better for predicting fMRI than the representations learned from fMRI alone, indicating that the learned representations capture brain-activity-relevant information that is not simply an artifact of the modality. While changes to language representations help the model predict brain activity, they also do not harm the model's ability to perform downstream NLP tasks. Our findings are notable for research on language understanding in the brain. Mariya Toneva, Leila Wehbe |
NeurIPS | 2 |
| 2019 | Interpreting and improving natural-language processing (in machines) with natural language-processing (in the brain)abstractNeural networks models for NLP are typically implemented without the explicit encoding of language rules and yet they are able to break one performance record after another. This has generated a lot of research interest in interpreting the representations learned by these networks. We propose here a novel interpretation approach that relies on the only processing system we have that does understand language: the human brain. We use brain imaging recordings of subjects reading complex natural text to interpret word and sequence embeddings from 4 recent NLP models - ELMo, USE, BERT and Transformer-XL. We study how their representations differ across layer depth, context length, and attention type. Our results reveal differences in the context-related representations across these models. Further, in the transformer models, we find an interaction between layer depth and context length, and between layer depth and attention type. We finally hypothesize that altering BERT to better align with brain recordings would enable it to also better understand language. Probing the altered BERT using syntactic NLP tasks reveals that the model with increased brain-alignment outperforms the original model. Cognitive neuroscientists have already begun using NLP networks to study the brain, and this work closes the loop to allow the interaction between NLP and cognitive neuroscience to be a true cross-pollination. Mariya Toneva, Leila Wehbe |
NeurIPS | 1 |
| 2014 | An Exploration of Social Grouping in Robots: Effects of Behavioral Mimicry, Appearance, and Eye Gaze
Ahsan Nawroj, Mariya Toneva, Henny Admoni, Brian Scassellati |
CogSci | 2 |
| 2012 | The Physical Presence of a Robot Tutor Increases Cognitive Learning Gains
Dan Leyzberg, Samuel Spaulding, Mariya Toneva, Brian Scassellati |
CogSci | 3 |
| 2011 | Robot gaze does not reflexively cue human attention
Henny Admoni, Caroline Bank, Mariya Toneva, Brian Scassellati |
CogSci | 4 |