Alexander G. Huth

dblp:44/8860 · also Alex Huth, Alexander Huth · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
8since 2021 · last 2024
0000-0002-5031-5348ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
11 papers
Representation and self-supervised learning · 29% Language models and text generation · 26% Deep learning architectures and training · 21%
Interdisciplinary, comprehensive, and emerging computing
6 papers
Bioinformatics and computational biology · 82% Medical and health informatics · 18%
Computer graphics and multimedia
2 papers
Visual content generation and editing · 80% Geometric modeling and processing · 20%
Theoretical computer science
2 papers
Algorithms and data structures · 60% Information theory · 40%

Topics — the 28 heaviest of 34, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
multimodal representation
1.222023
Scaling laws for language encoding models in fMRI · NeurIPS 2023
Low-dimensional Structure in the Space of Language Representations is Reflected in Brain Responses · NeurIPS 2021
Machine learning › Deep learning architectures and training
recurrent neural network
0.922021
Multi-timescale Representation Learning in LSTM Language Models · ICLR 2021
Approximating Stacked and Bidirectional Recurrent Architectures with the Delayed Recurrent Neural Network · ICML 2020
Natural language and speech › Language models and text generation › language modeling
LSTM language model
0.822021
Multi-timescale Representation Learning in LSTM Language Models · ICLR 2021
Incorporating Context into Language Encoding Models for fMRI · NeurIPS 2018
Machine learning › Representation and self-supervised learning › representation learning › embedding learning
interpretable embedding
0.812024
Crafting Interpretable Embeddings for Language Neuroscience by Asking LLMs Questions · NeurIPS 2024
Bioinformatics and computational biology
computational neuroscience
0.732024
Interpretable multi-timescale models for predicting fMRI responses to continuous natural speech · NeurIPS 2020
Crafting Interpretable Embeddings for Language Neuroscience by Asking LLMs Questions · NeurIPS 2024
Anthropic correction of information estimates and its application to neural coding · IEEE Trans. Inf. Theory 2010
Machine learning › Representation and self-supervised learning › computational neuroscience › neural coding
brain encoding models
0.712023
Scaling laws for language encoding models in fMRI · NeurIPS 2023
Natural language and speech › Language models and text generation › language model analysis
language model scaling
0.712023
Scaling laws for language encoding models in fMRI · NeurIPS 2023
Machine learning › Representation and self-supervised learning
multimodal representation learning
0.712023
Brain encoding models based on multimodal transformers can transfer across language and vision · NeurIPS 2023
Machine learning › Deep learning architectures and training
scaling laws
0.712023
Scaling laws for language encoding models in fMRI · NeurIPS 2023
Bioinformatics and computational biology › computational neuroscience › neural response modeling
brain encoding model
0.712023
Brain encoding models based on multimodal transformers can transfer across language and vision · NeurIPS 2023
Natural language and speech › Speech recognition and synthesis › speech representation learning
self-supervised speech representation
0.612022
Self-Supervised Models of Audio Effectively Explain Human Cortical Responses to Speech · ICML 2022
Machine learning › Efficient and distributed learning › data selection
data selection for fine-tuning
0.512021
Selecting Informative Contexts Improves Language Model Fine-tuning · ACL/IJCNLP (1) 2021
Natural language and speech › Language models and text generation
large language model fine-tuning
0.512021
Selecting Informative Contexts Improves Language Model Fine-tuning · ACL/IJCNLP (1) 2021
Machine learning › Representation and self-supervised learning
representation analysis
0.512021
Low-dimensional Structure in the Space of Language Representations is Reflected in Brain Responses · NeurIPS 2021
Machine learning › Deep learning architectures and training › recurrent neural network
bidirectional recurrent network
0.412020
Approximating Stacked and Bidirectional Recurrent Architectures with the Delayed Recurrent Neural Network · ICML 2020
Natural language and speech › Language models and text generation
neural language model
0.412020
Interpretable multi-timescale models for predicting fMRI responses to continuous natural speech · NeurIPS 2020
Visual content generation and editing
3d content generation
0.412020
Deep Generative Modeling for Scene Synthesis via Hybrid Representations · ACM Trans. Graph. 2020
Visual content generation and editing › 3d scene generation
indoor scene synthesis
0.412020
Deep Generative Modeling for Scene Synthesis via Hybrid Representations · ACM Trans. Graph. 2020
Visual content generation and editing
scene synthesis
0.412020
Deep Generative Modeling for Scene Synthesis via Hybrid Representations · ACM Trans. Graph. 2020
Medical and health informatics
neuroimaging
0.312018
Incorporating Context into Language Encoding Models for fMRI · NeurIPS 2018
Geometric modeling and processing
shape analysis
0.312018
Efficient, Sparse Representation of Manifold Distance Matrices for Classical Scaling · CVPR 2018
Algorithms and data structures › numerical linear algebra › dimensionality reduction
multidimensional scaling
0.312018
Efficient, Sparse Representation of Manifold Distance Matrices for Classical Scaling · CVPR 2018
Machine learning › Deep learning architectures and training › transformer
multimodal transformer
0.212023
Brain encoding models based on multimodal transformers can transfer across language and vision · NeurIPS 2023
Bioinformatics and computational biology › neuroscience
neuroinformatics
0.212022
Self-Supervised Models of Audio Effectively Explain Human Cortical Responses to Speech · ICML 2022
Natural language and speech › Language models and text generation › neural language model
neural language model representations
0.112021
Low-dimensional Structure in the Space of Language Representations is Reflected in Brain Responses · NeurIPS 2021
Information theory › estimation theory
bias correction
0.112010
Anthropic correction of information estimates and its application to neural coding · IEEE Trans. Inf. Theory 2010
Information theory › information measures › mutual information
mutual information estimation
0.112010
Anthropic correction of information estimates and its application to neural coding · IEEE Trans. Inf. Theory 2010
Bioinformatics and computational biology › computational neuroscience
neural coding
0.012010
Anthropic correction of information estimates and its application to neural coding · IEEE Trans. Inf. Theory 2010

Methods — techniques the papers use, named apart from their topics

prompting · 1.5large language model · 1.5transformer · 1.3multimodal pretraining · 1.3fMRI encoding · 1.2self-supervised learning · 1.1fMRI encoding models · 1.1transformer language model · 0.7sparse interpolation · 0.7noise ceiling analysis · 0.7biharmonic interpolation · 0.7context selection · 0.5encoding models · 0.4discriminative loss · 0.4LSTM · 0.43d object arrangement representation · 0.42d image representation · 0.4word embeddings · 0.3
YearPublicationVenuePosition
2024 Crafting Interpretable Embeddings for Language Neuroscience by Asking LLMs Questions
abstract
Large language models (LLMs) have rapidly improved text embeddings for a growing array of natural-language processing tasks. However, their opaqueness and proliferation into scientific domains such as neuroscience have created a growing need for interpretability. Here, we ask whether we can obtain interpretable embeddings through LLM prompting. We introduce question-answering embeddings (QA-Emb), embeddings where each feature represents an answer to a yes/no question asked to an LLM. Training QA-Emb reduces to selecting a set of underlying questions rather than learning model weights. We use QA-Emb to flexibly generate interpretable models for predicting fMRI voxel responses to language stimuli. QA-Emb significantly outperforms an established interpretable baseline, and does so while requiring very few questions. This paves the way towards building flexible feature spaces that can concretize and evaluate our understanding of semantic brain representations. We additionally find that QA-Emb can be effectively approximated with an efficient model, and we explore broader applications in simple NLP tasks.
Vinamra Benara, Chandan Singh, John X. Morris, Richard J. Antonello, Ion Stoica, Alexander G. Huth, Jianfeng Gao 0001
NeurIPS6
2023 Humans and language models diverge when predicting repeating text
abstract
Language models that are trained on the nextword prediction task have been shown to accurately model human behavior in word prediction and reading speed.In contrast with these findings, we present a scenario in which the performance of humans and LMs diverges.We collected a dataset of human next-word predictions for five stimuli that are formed by repeating spans of text.Human and GPT-2 LM predictions are strongly aligned in the first presentation of a text span, but their performance quickly diverges when memory (or in-context learning) begins to play a role.We traced the cause of this divergence to specific attention heads in a middle layer.Adding a power-law recency bias to these attention heads yielded a model that performs much more similarly to humans.We hope that this scenario will spur future work in bringing LMs closer to human behavior.1
Aditya R. Vaidya, Javier Turek, Alexander G. Huth
CoNLL3
2023 Scaling laws for language encoding models in fMRI
abstract
Representations from transformer-based unidirectional language models are known to be effective at predicting brain responses to natural language. However, most studies comparing language models to brains have used GPT-2 or similarly sized language models. Here we tested whether larger open-source models such as those from the OPT and LLaMA families are better at predicting brain responses recorded using fMRI. Mirroring scaling results from other contexts, we found that brain prediction performance scales logarithmically with model size from 125M to 30B parameter models, with ~15% increased encoding performance as measured by correlation with a held-out test set across 3 subjects. Similar log-linear behavior was observed when scaling the size of the fMRI training set. We also characterized scaling for acoustic encoding models that use HuBERT, WavLM, and Whisper, and we found comparable improvements with model size. A noise ceiling analysis of these large, high-performance encoding models showed that performance is nearing the theoretical maximum for brain areas such as the precuneus and higher auditory cortex. These results suggest that increasing scale in both models and data will yield incredibly effective models of language processing in the brain, enabling better scientific understanding as well as applications such as decoding.
Richard J. Antonello, Aditya R. Vaidya, Alexander G. Huth
NeurIPS3
2023 Brain encoding models based on multimodal transformers can transfer across language and vision
abstract
Encoding models have been used to assess how the human brain represents concepts in language and vision. While language and vision rely on similar concept representations, current encoding models are typically trained and tested on brain responses to each modality in isolation. Recent advances in multimodal pretraining have produced transformers that can extract aligned representations of concepts in language and vision. In this work, we used representations from multimodal transformers to train encoding models that can transfer across fMRI responses to stories and movies. We found that encoding models trained on brain responses to one modality can successfully predict brain responses to the other modality, particularly in cortical regions that represent conceptual meaning. Further analysis of these encoding models revealed shared semantic dimensions that underlie concept representations in language and vision. Comparing encoding models trained using representations from multimodal and unimodal transformers, we found that multimodal transformers learn more aligned representations of concepts in language and vision. Our results demonstrate how multimodal transformers can provide insights into the brain’s capacity for multimodal processing.
Jerry Tang, Vy A. Vo, Vasudev Lal, Alexander G. Huth
NeurIPS5
2022 Self-Supervised Models of Audio Effectively Explain Human Cortical Responses to Speech
abstract
Self-supervised language models are very effective at predicting high-level cortical responses during language comprehension. However, the best current models of lower-level auditory processing in the human brain rely on either hand-constructed acoustic filters or representations from supervised audio neural networks. In this work, we capitalize on the progress of self-supervised speech representation learning (SSL) to create new state-of-the-art models of the human auditory system. Compared against acoustic baselines, phonemic features, and supervised models, representations from the middle layers of self-supervised models (APC, wav2vec, wav2vec 2.0, and HuBERT) consistently yield the best prediction performance for fMRI recordings within the auditory cortex (AC). Brain areas involved in low-level auditory processing exhibit a preference for earlier SSL model layers, whereas higher-level semantic areas prefer later layers. We show that these trends are due to the models’ ability to encode information at multiple linguistic levels (acoustic, phonetic, and lexical) along their representation depth. Overall, these results show that self-supervised models effectively capture the hierarchy of information relevant to different stages of speech processing in human cortex.
Aditya R. Vaidya, Shailee Jain, Alexander G. Huth
ICML3
2021 Selecting Informative Contexts Improves Language Model Fine-tuning
abstract
Richard Antonello, Nicole Beckage, Javier Turek, Alexander Huth. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Richard J. Antonello, Nicole Beckage, Javier Turek, Alexander G. Huth
ACL/IJCNLP (1)4
2021 Multi-timescale Representation Learning in LSTM Language Models
Shivangi Mahto, Vy A. Vo, Javier Turek, Alexander G. Huth
ICLR4
2021 Low-dimensional Structure in the Space of Language Representations is Reflected in Brain Responses
abstract
How related are the representations learned by neural language models, translation models, and language tagging tasks? We answer this question by adapting an encoder-decoder transfer learning method from computer vision to investigate the structure among 100 different feature spaces extracted from hidden representations of various networks trained on language tasks.This method reveals a low-dimensional structure where language models and translation models smoothly interpolate between word embeddings, syntactic and semantic tasks, and future word embeddings. We call this low-dimensional structure a language representation embedding because it encodes the relationships between representations needed to process language for a variety of NLP tasks. We find that this representation embedding can predict how well each individual feature space maps to human brain responses to natural language stimuli recorded using fMRI. Additionally, we find that the principal dimension of this structure can be used to create a metric which highlights the brain's natural language processing hierarchy. This suggests that the embedding captures some part of the brain's natural language representation structure.
Richard J. Antonello, Javier Turek, Vy A. Vo, Alexander G. Huth
NeurIPS4
2020 Approximating Stacked and Bidirectional Recurrent Architectures with the Delayed Recurrent Neural Network
abstract
Recent work has shown that topological enhancements to recurrent neural networks (RNNs) can increase their expressiveness and representational capacity. Two popular enhancements are stacked RNNs, which increases the capacity for learning non-linear functions, and bidirectional processing, which exploits acausal information in a sequence. In this work, we explore the delayed-RNN, which is a single-layer RNN that has a delay between the input and output. We prove that a weight-constrained version of the delayed-RNN is equivalent to a stacked-RNN. We also show that the delay gives rise to partial acausality, much like bidirectional networks. Synthetic experiments confirm that the delayed-RNN can mimic bidirectional networks, solving some acausal tasks similarly, and outperforming them in others. Moreover, we show similar performance to bidirectional networks in a real-world natural language processing task. These results suggest that delayed-RNNs can approximate topologies including stacked RNNs, bidirectional RNNs, and stacked bidirectional RNNs – but with equivalent or faster runtimes for the delayed-RNNs.
Javier Turek, Shailee Jain, Vy A. Vo, Mihai Capota, Alexander G. Huth, Theodore L. Willke
ICML5
2020 Interpretable multi-timescale models for predicting fMRI responses to continuous natural speech
abstract
Natural language contains information at multiple timescales. To understand how the human brain represents this information, one approach is to build encoding models that predict fMRI responses to natural language using representations extracted from neural network language models (LMs). However, these LM-derived representations do not explicitly separate information at different timescales, making it difficult to interpret the encoding models. In this work we construct interpretable multi-timescale representations by forcing individual units in an LSTM LM to integrate information over specific temporal scales. This allows us to explicitly and directly map the timescale of information encoded by each individual fMRI voxel. Further, the standard fMRI encoding procedure does not account for varying temporal properties in the encoding features. We modify the procedure so that it can capture both short- and long-timescale information. This approach outperforms other encoding models, particularly for voxels that represent long-timescale information. It also provides a finer-grained map of timescale information in the human language pathway. This serves as a framework for future work investigating temporal hierarchies across artificial and biological language systems.
Shailee Jain, Vy A. Vo, Shivangi Mahto, Amanda LeBel, Javier Turek, Alexander G. Huth
NeurIPS6
2020 Deep Generative Modeling for Scene Synthesis via Hybrid Representations
abstract
We present a deep generative scene modeling technique for indoor environments. Our goal is to train a generative model using a feed-forward neural network that maps a prior distribution (e.g., a normal distribution) to the distribution of primary objects in indoor scenes. We introduce a 3D object arrangement representation that models the locations and orientations of objects, based on their size and shape attributes. Moreover, our scene representation is applicable for 3D objects with different multiplicities (repetition counts), selected from a database. We show a principled way to train this model by combining discriminative losses for both a 3D object arrangement representation and a 2D image-based representation. We demonstrate the effectiveness of our scene representation and the network training method on benchmark datasets. We also show the applications of this generative model in scene interpolation and scene completion.
Zaiwei Zhang, Zhenpei Yang, Chongyang Ma, Linjie Luo, Alexander G. Huth, Etienne Vouga, Qixing Huang
ACM Trans. Graph.5
2018 Efficient, Sparse Representation of Manifold Distance Matrices for Classical Scaling
abstract
Geodesic distance matrices can reveal shape properties that are largely invariant to non-rigid deformations, and thus are often used to analyze and represent 3-D shapes. However, these matrices grow quadratically with the number of points. Thus for large point sets it is common to use a low-rank approximation to the distance matrix, which fits in memory and can be efficiently analyzed using methods such as multidimensional scaling (MDS). In this paper we present a novel sparse method for efficiently representing geodesic distance matrices using biharmonic interpolation. This method exploits knowledge of the data manifold to learn a sparse interpolation operator that approximates distances using a subset of points. We show that our method is 2x faster and uses 20x less memory than current leading methods for solving MDS on large point sets, with similar quality. This enables analyses of large point sets that were previously infeasible.
Javier Turek, Alexander G. Huth
CVPR2
2018 Incorporating Context into Language Encoding Models for fMRI
abstract
Language encoding models help explain language processing in the human brain by learning functions that predict brain responses from the language stimuli that elicited them. Current word embedding-based approaches treat each stimulus word independently and thus ignore the influence of context on language understanding. In this work we instead build encoding models using rich contextual representations derived from an LSTM language model. Our models show a significant improvement in encoding performance relative to state-of-the-art embeddings in nearly every brain area. By varying the amount of context used in the models and providing the models with distorted context, we show that this improvement is due to a combination of better word embeddings learned by the LSTM language model and contextual information. We are also able to use our models to map context sensitivity across the cortex. These results suggest that LSTM language models learn high-level representations that are related to representations in the human brain.
Shailee Jain, Alexander G. Huth
NeurIPS2
2010 Anthropic correction of information estimates and its application to neural coding
abstract
Information theory has been used as an organizing principle in neuroscience for several decades. Estimates of the mutual information (MI) between signals acquired in neurophysiological experiments are believed to yield insights into the structure of the underlying information processing architectures. With the pervasive availability of recordings from many neurons, several information and redundancy measures have been proposed in the recent literature. A typical scenario is that only a small number of stimuli can be tested, while ample response data may be available for each of the tested stimuli. The resulting asymmetric information estimation problem is considered. It is shown that the direct plug-in information estimate has a negative bias. An anthropic correction is introduced that has a positive bias. These two complementary estimators and their combinations are natural candidates for information estimation in neuroscience. Tail and variance bounds are given for both estimates. The proposed information estimates are applied to the analysis of neural discrimination and redundancy in the avian auditory system.
Michael Gastpar, Patrick R. Gill, Alexander G. Huth, Frédéric E. Theunissen
IEEE Trans. Inf. Theory3