Richard J. Antonello

dblp:387/2006 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Representation and self-supervised learning · 40% Language models and text generation · 19% Vision and language · 17%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 100%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
multimodal representation
1.222023
Scaling laws for language encoding models in fMRI · NeurIPS 2023
Low-dimensional Structure in the Space of Language Representations is Reflected in Brain Responses · NeurIPS 2021
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning
0.912025
Far from the Shallow: Brain-Predictive Reasoning Embedding through Residual Disentanglement · NeurIPS 2025
Bioinformatics and computational biology › computational neuroscience › neural response modeling
brain encoding model
0.912025
Far from the Shallow: Brain-Predictive Reasoning Embedding through Residual Disentanglement · NeurIPS 2025
Bioinformatics and computational biology › neuroscience
neuroinformatics
0.912025
Far from the Shallow: Brain-Predictive Reasoning Embedding through Residual Disentanglement · NeurIPS 2025
Machine learning › Representation and self-supervised learning › representation learning › embedding learning
interpretable embedding
0.812024
Crafting Interpretable Embeddings for Language Neuroscience by Asking LLMs Questions · NeurIPS 2024
Machine learning › Representation and self-supervised learning › computational neuroscience › neural coding
brain encoding models
0.712023
Scaling laws for language encoding models in fMRI · NeurIPS 2023
Natural language and speech › Language models and text generation › language model analysis
language model scaling
0.712023
Scaling laws for language encoding models in fMRI · NeurIPS 2023
Machine learning › Deep learning architectures and training
scaling laws
0.712023
Scaling laws for language encoding models in fMRI · NeurIPS 2023
Machine learning › Efficient and distributed learning › data selection
data selection for fine-tuning
0.512021
Selecting Informative Contexts Improves Language Model Fine-tuning · ACL/IJCNLP (1) 2021
Natural language and speech › Language models and text generation
large language model fine-tuning
0.512021
Selecting Informative Contexts Improves Language Model Fine-tuning · ACL/IJCNLP (1) 2021
Machine learning › Representation and self-supervised learning
representation analysis
0.512021
Low-dimensional Structure in the Space of Language Representations is Reflected in Brain Responses · NeurIPS 2021
Machine learning › Trustworthy machine learning › language model interpretability
language model probing
0.312025
Far from the Shallow: Brain-Predictive Reasoning Embedding through Residual Disentanglement · NeurIPS 2025
Bioinformatics and computational biology
computational neuroscience
0.212024
Crafting Interpretable Embeddings for Language Neuroscience by Asking LLMs Questions · NeurIPS 2024
Natural language and speech › Language models and text generation › neural language model
neural language model representations
0.112021
Low-dimensional Structure in the Space of Language Representations is Reflected in Brain Responses · NeurIPS 2021

Methods — techniques the papers use, named apart from their topics

residual disentanglement · 1.7regression · 1.7probing · 1.7prompting · 1.5large language model · 1.5fMRI encoding · 1.2transformer language model · 0.7noise ceiling analysis · 0.7fine-tuning · 0.5context selection · 0.5
YearPublicationVenuePosition
2025 Neuro2Semantic: A Transfer Learning Framework for Semantic Reconstruction of Continuous Language from Human Intracranial EEG
Siavash Shams, Richard J. Antonello, Gavin Mischler, Stephan Bickel, Ashesh D. Mehta, Nima Mesgarani
INTERSPEECH2
2025 Far from the Shallow: Brain-Predictive Reasoning Embedding through Residual Disentanglement
abstract
Understanding how the human brain progresses from processing simple linguistic inputs to performing high-level reasoning is a fundamental challenge in neuroscience. While modern large language models (LLMs) are increasingly used to model neural responses to language, their internal representations are highly "entangled," mixing information about lexicon, syntax, meaning, and reasoning. This entanglement biases conventional brain encoding analyses toward linguistically shallow features (e.g., lexicon and syntax), making it difficult to isolate the neural substrates of cognitively deeper processes. Here, we introduce a residual disentanglement method that computationally isolates these components. By first probing an LM to identify feature-specific layers, our method iteratively regresses out lower-level representations to produce four nearly orthogonal embeddings for lexicon, syntax, meaning, and, critically, reasoning. We used these disentangled embeddings to model intracranial (ECoG) brain recordings from neurosurgical patients listening to natural speech. We show that: 1) This isolated reasoning embedding exhibits unique predictive power, accounting for variance in neural activity not explained by other linguistic features and even extending to the recruitment of visual regions beyond classical language areas. 2) The neural signature for reasoning is temporally distinct, peaking later (~350-400ms) than signals related to lexicon, syntax, and meaning, consistent with its position atop a processing hierarchy. 3) Standard, non-disentangled LLM embeddings can be misleading, as their predictive success is primarily attributable to linguistically shallow features, masking the more subtle contributions of deeper cognitive processing. Our work provides compelling neural evidence for an abstract reasoning computation during language comprehension and offers a robust framework for mapping distinct cognitive functions from artificial models to the human brain.
Linyang He, Tianjun Zhong, Richard J. Antonello, Gavin Mischler, Micah Goldblum, Nima Mesgarani
NeurIPS3
2024 Crafting Interpretable Embeddings for Language Neuroscience by Asking LLMs Questions
abstract
Large language models (LLMs) have rapidly improved text embeddings for a growing array of natural-language processing tasks. However, their opaqueness and proliferation into scientific domains such as neuroscience have created a growing need for interpretability. Here, we ask whether we can obtain interpretable embeddings through LLM prompting. We introduce question-answering embeddings (QA-Emb), embeddings where each feature represents an answer to a yes/no question asked to an LLM. Training QA-Emb reduces to selecting a set of underlying questions rather than learning model weights. We use QA-Emb to flexibly generate interpretable models for predicting fMRI voxel responses to language stimuli. QA-Emb significantly outperforms an established interpretable baseline, and does so while requiring very few questions. This paves the way towards building flexible feature spaces that can concretize and evaluate our understanding of semantic brain representations. We additionally find that QA-Emb can be effectively approximated with an efficient model, and we explore broader applications in simple NLP tasks.
Vinamra Benara, Chandan Singh, John X. Morris, Richard J. Antonello, Ion Stoica, Alexander G. Huth, Jianfeng Gao 0001
NeurIPS4
2023 Scaling laws for language encoding models in fMRI
abstract
Representations from transformer-based unidirectional language models are known to be effective at predicting brain responses to natural language. However, most studies comparing language models to brains have used GPT-2 or similarly sized language models. Here we tested whether larger open-source models such as those from the OPT and LLaMA families are better at predicting brain responses recorded using fMRI. Mirroring scaling results from other contexts, we found that brain prediction performance scales logarithmically with model size from 125M to 30B parameter models, with ~15% increased encoding performance as measured by correlation with a held-out test set across 3 subjects. Similar log-linear behavior was observed when scaling the size of the fMRI training set. We also characterized scaling for acoustic encoding models that use HuBERT, WavLM, and Whisper, and we found comparable improvements with model size. A noise ceiling analysis of these large, high-performance encoding models showed that performance is nearing the theoretical maximum for brain areas such as the precuneus and higher auditory cortex. These results suggest that increasing scale in both models and data will yield incredibly effective models of language processing in the brain, enabling better scientific understanding as well as applications such as decoding.
Richard J. Antonello, Aditya R. Vaidya, Alexander G. Huth
NeurIPS1
2021 Selecting Informative Contexts Improves Language Model Fine-tuning
abstract
Richard Antonello, Nicole Beckage, Javier Turek, Alexander Huth. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Richard J. Antonello, Nicole Beckage, Javier Turek, Alexander G. Huth
ACL/IJCNLP (1)1
2021 Low-dimensional Structure in the Space of Language Representations is Reflected in Brain Responses
abstract
How related are the representations learned by neural language models, translation models, and language tagging tasks? We answer this question by adapting an encoder-decoder transfer learning method from computer vision to investigate the structure among 100 different feature spaces extracted from hidden representations of various networks trained on language tasks.This method reveals a low-dimensional structure where language models and translation models smoothly interpolate between word embeddings, syntactic and semantic tasks, and future word embeddings. We call this low-dimensional structure a language representation embedding because it encodes the relationships between representations needed to process language for a variety of NLP tasks. We find that this representation embedding can predict how well each individual feature space maps to human brain responses to natural language stimuli recorded using fMRI. Additionally, we find that the principal dimension of this structure can be used to create a metric which highlights the brain's natural language processing hierarchy. This suggests that the embedding captures some part of the brain's natural language representation structure.
Richard J. Antonello, Javier Turek, Vy A. Vo, Alexander G. Huth
NeurIPS1