EDBT 2026 Demo / reviewers in the wild / expert
Alexander G. Huth
dblp:44/8860 · also Alex Huth, Alexander Huth
· DBLP profile ↗
14ranked-venue papers
0as first author
8since 2021 · last 2024
0000-0002-5031-5348ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Theory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
11 papers |
Representation and self-supervised learning · 29% Language models and text generation · 26% Deep learning architectures and training · 21% | |
| Interdisciplinary, comprehensive, and emerging computing
6 papers |
Bioinformatics and computational biology · 82% Medical and health informatics · 18% | |
| Computer graphics and multimedia
2 papers |
Visual content generation and editing · 80% Geometric modeling and processing · 20% | |
| Theoretical computer science
2 papers |
Algorithms and data structures · 60% Information theory · 40% |
Topics — the 28 heaviest of 34, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
multimodal representation |
1.2 | 2 | 2023 | Scaling laws for language encoding models in fMRI · NeurIPS 2023 Low-dimensional Structure in the Space of Language Representations is Reflected in Brain Responses · NeurIPS 2021 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.9 | 2 | 2021 | Multi-timescale Representation Learning in LSTM Language Models · ICLR 2021 Approximating Stacked and Bidirectional Recurrent Architectures with the Delayed Recurrent Neural Network · ICML 2020 |
Natural language and speech › Language models and text generation › language modeling
LSTM language model |
0.8 | 2 | 2021 | Multi-timescale Representation Learning in LSTM Language Models · ICLR 2021 Incorporating Context into Language Encoding Models for fMRI · NeurIPS 2018 |
Machine learning › Representation and self-supervised learning › representation learning › embedding learning
interpretable embedding |
0.8 | 1 | 2024 | Crafting Interpretable Embeddings for Language Neuroscience by Asking LLMs Questions · NeurIPS 2024 |
Bioinformatics and computational biology
computational neuroscience |
0.7 | 3 | 2024 | Interpretable multi-timescale models for predicting fMRI responses to continuous natural speech · NeurIPS 2020 Crafting Interpretable Embeddings for Language Neuroscience by Asking LLMs Questions · NeurIPS 2024 Anthropic correction of information estimates and its application to neural coding · IEEE Trans. Inf. Theory 2010 |
Machine learning › Representation and self-supervised learning › computational neuroscience › neural coding
brain encoding models |
0.7 | 1 | 2023 | Scaling laws for language encoding models in fMRI · NeurIPS 2023 |
Natural language and speech › Language models and text generation › language model analysis
language model scaling |
0.7 | 1 | 2023 | Scaling laws for language encoding models in fMRI · NeurIPS 2023 |
Machine learning › Representation and self-supervised learning
multimodal representation learning |
0.7 | 1 | 2023 | Brain encoding models based on multimodal transformers can transfer across language and vision · NeurIPS 2023 |
Machine learning › Deep learning architectures and training
scaling laws |
0.7 | 1 | 2023 | Scaling laws for language encoding models in fMRI · NeurIPS 2023 |
Bioinformatics and computational biology › computational neuroscience › neural response modeling
brain encoding model |
0.7 | 1 | 2023 | Brain encoding models based on multimodal transformers can transfer across language and vision · NeurIPS 2023 |
Natural language and speech › Speech recognition and synthesis › speech representation learning
self-supervised speech representation |
0.6 | 1 | 2022 | Self-Supervised Models of Audio Effectively Explain Human Cortical Responses to Speech · ICML 2022 |
Machine learning › Efficient and distributed learning › data selection
data selection for fine-tuning |
0.5 | 1 | 2021 | Selecting Informative Contexts Improves Language Model Fine-tuning · ACL/IJCNLP (1) 2021 |
Natural language and speech › Language models and text generation
large language model fine-tuning |
0.5 | 1 | 2021 | Selecting Informative Contexts Improves Language Model Fine-tuning · ACL/IJCNLP (1) 2021 |
Machine learning › Representation and self-supervised learning
representation analysis |
0.5 | 1 | 2021 | Low-dimensional Structure in the Space of Language Representations is Reflected in Brain Responses · NeurIPS 2021 |
Machine learning › Deep learning architectures and training › recurrent neural network
bidirectional recurrent network |
0.4 | 1 | 2020 | Approximating Stacked and Bidirectional Recurrent Architectures with the Delayed Recurrent Neural Network · ICML 2020 |
Natural language and speech › Language models and text generation
neural language model |
0.4 | 1 | 2020 | Interpretable multi-timescale models for predicting fMRI responses to continuous natural speech · NeurIPS 2020 |
Visual content generation and editing
3d content generation |
0.4 | 1 | 2020 | Deep Generative Modeling for Scene Synthesis via Hybrid Representations · ACM Trans. Graph. 2020 |
Visual content generation and editing › 3d scene generation
indoor scene synthesis |
0.4 | 1 | 2020 | Deep Generative Modeling for Scene Synthesis via Hybrid Representations · ACM Trans. Graph. 2020 |
Visual content generation and editing
scene synthesis |
0.4 | 1 | 2020 | Deep Generative Modeling for Scene Synthesis via Hybrid Representations · ACM Trans. Graph. 2020 |
Medical and health informatics
neuroimaging |
0.3 | 1 | 2018 | Incorporating Context into Language Encoding Models for fMRI · NeurIPS 2018 |
Geometric modeling and processing
shape analysis |
0.3 | 1 | 2018 | Efficient, Sparse Representation of Manifold Distance Matrices for Classical Scaling · CVPR 2018 |
Algorithms and data structures › numerical linear algebra › dimensionality reduction
multidimensional scaling |
0.3 | 1 | 2018 | Efficient, Sparse Representation of Manifold Distance Matrices for Classical Scaling · CVPR 2018 |
Machine learning › Deep learning architectures and training › transformer
multimodal transformer |
0.2 | 1 | 2023 | Brain encoding models based on multimodal transformers can transfer across language and vision · NeurIPS 2023 |
Bioinformatics and computational biology › neuroscience
neuroinformatics |
0.2 | 1 | 2022 | Self-Supervised Models of Audio Effectively Explain Human Cortical Responses to Speech · ICML 2022 |
Natural language and speech › Language models and text generation › neural language model
neural language model representations |
0.1 | 1 | 2021 | Low-dimensional Structure in the Space of Language Representations is Reflected in Brain Responses · NeurIPS 2021 |
Information theory › estimation theory
bias correction |
0.1 | 1 | 2010 | Anthropic correction of information estimates and its application to neural coding · IEEE Trans. Inf. Theory 2010 |
Information theory › information measures › mutual information
mutual information estimation |
0.1 | 1 | 2010 | Anthropic correction of information estimates and its application to neural coding · IEEE Trans. Inf. Theory 2010 |
Bioinformatics and computational biology › computational neuroscience
neural coding |
0.0 | 1 | 2010 | Anthropic correction of information estimates and its application to neural coding · IEEE Trans. Inf. Theory 2010 |
Methods — techniques the papers use, named apart from their topics
prompting · 1.5large language model · 1.5transformer · 1.3multimodal pretraining · 1.3fMRI encoding · 1.2self-supervised learning · 1.1fMRI encoding models · 1.1transformer language model · 0.7sparse interpolation · 0.7noise ceiling analysis · 0.7biharmonic interpolation · 0.7context selection · 0.5encoding models · 0.4discriminative loss · 0.4LSTM · 0.43d object arrangement representation · 0.42d image representation · 0.4word embeddings · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Crafting Interpretable Embeddings for Language Neuroscience by Asking LLMs QuestionsabstractLarge language models (LLMs) have rapidly improved text embeddings for a growing array of natural-language processing tasks. However, their opaqueness and proliferation into scientific domains such as neuroscience have created a growing need for interpretability. Here, we ask whether we can obtain interpretable embeddings through LLM prompting. We introduce question-answering embeddings (QA-Emb), embeddings where each feature represents an answer to a yes/no question asked to an LLM. Training QA-Emb reduces to selecting a set of underlying questions rather than learning model weights.
We use QA-Emb to flexibly generate interpretable models for predicting fMRI voxel responses to language stimuli. QA-Emb significantly outperforms an established interpretable baseline, and does so while requiring very few questions. This paves the way towards building flexible feature spaces that can concretize and evaluate our understanding of semantic brain representations. We additionally find that QA-Emb can be effectively approximated with an efficient model, and we explore broader applications in simple NLP tasks. Vinamra Benara, Chandan Singh, John X. Morris, Richard J. Antonello, Ion Stoica, Alexander G. Huth, Jianfeng Gao 0001 |
NeurIPS | 6 |
| 2023 | Humans and language models diverge when predicting repeating textabstractLanguage models that are trained on the nextword prediction task have been shown to accurately model human behavior in word prediction and reading speed.In contrast with these findings, we present a scenario in which the performance of humans and LMs diverges.We collected a dataset of human next-word predictions for five stimuli that are formed by repeating spans of text.Human and GPT-2 LM predictions are strongly aligned in the first presentation of a text span, but their performance quickly diverges when memory (or in-context learning) begins to play a role.We traced the cause of this divergence to specific attention heads in a middle layer.Adding a power-law recency bias to these attention heads yielded a model that performs much more similarly to humans.We hope that this scenario will spur future work in bringing LMs closer to human behavior.1 Aditya R. Vaidya, Javier Turek, Alexander G. Huth |
CoNLL | 3 |
| 2023 | Scaling laws for language encoding models in fMRIabstractRepresentations from transformer-based unidirectional language models are known to be effective at predicting brain responses to natural language. However, most studies comparing language models to brains have used GPT-2 or similarly sized language models. Here we tested whether larger open-source models such as those from the OPT and LLaMA families are better at predicting brain responses recorded using fMRI. Mirroring scaling results from other contexts, we found that brain prediction performance scales logarithmically with model size from 125M to 30B parameter models, with ~15% increased encoding performance as measured by correlation with a held-out test set across 3 subjects. Similar log-linear behavior was observed when scaling the size of the fMRI training set. We also characterized scaling for acoustic encoding models that use HuBERT, WavLM, and Whisper, and we found comparable improvements with model size. A noise ceiling analysis of these large, high-performance encoding models showed that performance is nearing the theoretical maximum for brain areas such as the precuneus and higher auditory cortex. These results suggest that increasing scale in both models and data will yield incredibly effective models of language processing in the brain, enabling better scientific understanding as well as applications such as decoding. Richard J. Antonello, Aditya R. Vaidya, Alexander G. Huth |
NeurIPS | 3 |
| 2023 | Brain encoding models based on multimodal transformers can transfer across language and visionabstractEncoding models have been used to assess how the human brain represents concepts in language and vision. While language and vision rely on similar concept representations, current encoding models are typically trained and tested on brain responses to each modality in isolation. Recent advances in multimodal pretraining have produced transformers that can extract aligned representations of concepts in language and vision. In this work, we used representations from multimodal transformers to train encoding models that can transfer across fMRI responses to stories and movies. We found that encoding models trained on brain responses to one modality can successfully predict brain responses to the other modality, particularly in cortical regions that represent conceptual meaning. Further analysis of these encoding models revealed shared semantic dimensions that underlie concept representations in language and vision. Comparing encoding models trained using representations from multimodal and unimodal transformers, we found that multimodal transformers learn more aligned representations of concepts in language and vision. Our results demonstrate how multimodal transformers can provide insights into the brain’s capacity for multimodal processing. Jerry Tang, Vy A. Vo, Vasudev Lal, Alexander G. Huth |
NeurIPS | 5 |
| 2022 | Self-Supervised Models of Audio Effectively Explain Human Cortical Responses to SpeechabstractSelf-supervised language models are very effective at predicting high-level cortical responses during language comprehension. However, the best current models of lower-level auditory processing in the human brain rely on either hand-constructed acoustic filters or representations from supervised audio neural networks. In this work, we capitalize on the progress of self-supervised speech representation learning (SSL) to create new state-of-the-art models of the human auditory system. Compared against acoustic baselines, phonemic features, and supervised models, representations from the middle layers of self-supervised models (APC, wav2vec, wav2vec 2.0, and HuBERT) consistently yield the best prediction performance for fMRI recordings within the auditory cortex (AC). Brain areas involved in low-level auditory processing exhibit a preference for earlier SSL model layers, whereas higher-level semantic areas prefer later layers. We show that these trends are due to the models’ ability to encode information at multiple linguistic levels (acoustic, phonetic, and lexical) along their representation depth. Overall, these results show that self-supervised models effectively capture the hierarchy of information relevant to different stages of speech processing in human cortex. Aditya R. Vaidya, Shailee Jain, Alexander G. Huth |
ICML | 3 |
| 2021 | Selecting Informative Contexts Improves Language Model Fine-tuningabstractRichard Antonello, Nicole Beckage, Javier Turek, Alexander Huth. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Richard J. Antonello, Nicole Beckage, Javier Turek, Alexander G. Huth |
ACL/IJCNLP (1) | 4 |
| 2021 | Multi-timescale Representation Learning in LSTM Language Models
Shivangi Mahto, Vy A. Vo, Javier Turek, Alexander G. Huth |
ICLR | 4 |
| 2021 | Low-dimensional Structure in the Space of Language Representations is Reflected in Brain ResponsesabstractHow related are the representations learned by neural language models, translation models, and language tagging tasks? We answer this question by adapting an encoder-decoder transfer learning method from computer vision to investigate the structure among 100 different feature spaces extracted from hidden representations of various networks trained on language tasks.This method reveals a low-dimensional structure where language models and translation models smoothly interpolate between word embeddings, syntactic and semantic tasks, and future word embeddings. We call this low-dimensional structure a language representation embedding because it encodes the relationships between representations needed to process language for a variety of NLP tasks. We find that this representation embedding can predict how well each individual feature space maps to human brain responses to natural language stimuli recorded using fMRI. Additionally, we find that the principal dimension of this structure can be used to create a metric which highlights the brain's natural language processing hierarchy. This suggests that the embedding captures some part of the brain's natural language representation structure. Richard J. Antonello, Javier Turek, Vy A. Vo, Alexander G. Huth |
NeurIPS | 4 |
| 2020 | Approximating Stacked and Bidirectional Recurrent Architectures with the Delayed Recurrent Neural NetworkabstractRecent work has shown that topological enhancements to recurrent neural networks (RNNs) can increase their expressiveness and representational capacity. Two popular enhancements are stacked RNNs, which increases the capacity for learning non-linear functions, and bidirectional processing, which exploits acausal information in a sequence. In this work, we explore the delayed-RNN, which is a single-layer RNN that has a delay between the input and output. We prove that a weight-constrained version of the delayed-RNN is equivalent to a stacked-RNN. We also show that the delay gives rise to partial acausality, much like bidirectional networks. Synthetic experiments confirm that the delayed-RNN can mimic bidirectional networks, solving some acausal tasks similarly, and outperforming them in others. Moreover, we show similar performance to bidirectional networks in a real-world natural language processing task. These results suggest that delayed-RNNs can approximate topologies including stacked RNNs, bidirectional RNNs, and stacked bidirectional RNNs – but with equivalent or faster runtimes for the delayed-RNNs. Javier Turek, Shailee Jain, Vy A. Vo, Mihai Capota, Alexander G. Huth, Theodore L. Willke |
ICML | 5 |
| 2020 | Interpretable multi-timescale models for predicting fMRI responses to continuous natural speechabstractNatural language contains information at multiple timescales. To understand how the human brain represents this information, one approach is to build encoding models that predict fMRI responses to natural language using representations extracted from neural network language models (LMs). However, these LM-derived representations do not explicitly separate information at different timescales, making it difficult to interpret the encoding models. In this work we construct interpretable multi-timescale representations by forcing individual units in an LSTM LM to integrate information over specific temporal scales. This allows us to explicitly and directly map the timescale of information encoded by each individual fMRI voxel. Further, the standard fMRI encoding procedure does not account for varying temporal properties in the encoding features. We modify the procedure so that it can capture both short- and long-timescale information. This approach outperforms other encoding models, particularly for voxels that represent long-timescale information. It also provides a finer-grained map of timescale information in the human language pathway. This serves as a framework for future work investigating temporal hierarchies across artificial and biological language systems. Shailee Jain, Vy A. Vo, Shivangi Mahto, Amanda LeBel, Javier Turek, Alexander G. Huth |
NeurIPS | 6 |
| 2020 | Deep Generative Modeling for Scene Synthesis via Hybrid RepresentationsabstractWe present a deep generative scene modeling technique for indoor environments. Our goal is to train a generative model using a feed-forward neural network that maps a prior distribution (e.g., a normal distribution) to the distribution of primary objects in indoor scenes. We introduce a 3D object arrangement representation that models the locations and orientations of objects, based on their size and shape attributes. Moreover, our scene representation is applicable for 3D objects with different multiplicities (repetition counts), selected from a database. We show a principled way to train this model by combining discriminative losses for both a 3D object arrangement representation and a 2D image-based representation. We demonstrate the effectiveness of our scene representation and the network training method on benchmark datasets. We also show the applications of this generative model in scene interpolation and scene completion. Zaiwei Zhang, Zhenpei Yang, Chongyang Ma, Linjie Luo, Alexander G. Huth, Etienne Vouga, Qixing Huang |
ACM Trans. Graph. | 5 |
| 2018 | Efficient, Sparse Representation of Manifold Distance Matrices for Classical ScalingabstractGeodesic distance matrices can reveal shape properties that are largely invariant to non-rigid deformations, and thus are often used to analyze and represent 3-D shapes. However, these matrices grow quadratically with the number of points. Thus for large point sets it is common to use a low-rank approximation to the distance matrix, which fits in memory and can be efficiently analyzed using methods such as multidimensional scaling (MDS). In this paper we present a novel sparse method for efficiently representing geodesic distance matrices using biharmonic interpolation. This method exploits knowledge of the data manifold to learn a sparse interpolation operator that approximates distances using a subset of points. We show that our method is 2x faster and uses 20x less memory than current leading methods for solving MDS on large point sets, with similar quality. This enables analyses of large point sets that were previously infeasible. Javier Turek, Alexander G. Huth |
CVPR | 2 |
| 2018 | Incorporating Context into Language Encoding Models for fMRIabstractLanguage encoding models help explain language processing in the human brain by learning functions that predict brain responses from the language stimuli that elicited them. Current word embedding-based approaches treat each stimulus word independently and thus ignore the influence of context on language understanding. In this work we instead build encoding models using rich contextual representations derived from an LSTM language model. Our models show a significant improvement in encoding performance relative to state-of-the-art embeddings in nearly every brain area. By varying the amount of context used in the models and providing the models with distorted context, we show that this improvement is due to a combination of better word embeddings learned by the LSTM language model and contextual information. We are also able to use our models to map context sensitivity across the cortex. These results suggest that LSTM language models learn high-level representations that are related to representations in the human brain. Shailee Jain, Alexander G. Huth |
NeurIPS | 2 |
| 2010 | Anthropic correction of information estimates and its application to neural codingabstractInformation theory has been used as an organizing principle in neuroscience for several decades. Estimates of the mutual information (MI) between signals acquired in neurophysiological experiments are believed to yield insights into the structure of the underlying information processing architectures. With the pervasive availability of recordings from many neurons, several information and redundancy measures have been proposed in the recent literature. A typical scenario is that only a small number of stimuli can be tested, while ample response data may be available for each of the tested stimuli. The resulting asymmetric information estimation problem is considered. It is shown that the direct plug-in information estimate has a negative bias. An anthropic correction is introduced that has a positive bias. These two complementary estimators and their combinations are natural candidates for information estimation in neuroscience. Tail and variance bounds are given for both estimates. The proposed information estimates are applied to the analysis of neural discrimination and redundancy in the avian auditory system. Michael Gastpar, Patrick R. Gill, Alexander G. Huth, Frédéric E. Theunissen |
IEEE Trans. Inf. Theory | 3 |