Armin W. Thomas

dblp:228/8292 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2025
0000-0002-9947-5705ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Deep learning architectures and training · 52% Efficient and distributed learning · 36% Language models and text generation · 13%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 88% Medical and health informatics · 12%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
sequence modeling
2.232025
Quantifying Memory Utilization with Effective State-Size · ICML 2025
Simple Hardware-Efficient Long Convolutions for Sequence Modeling · ICML 2023
Hungry Hungry Hippos: Towards Language Modeling with State Space Models · ICLR 2023
Machine learning › Deep learning architectures and training
state space model
1.732025
Quantifying Memory Utilization with Effective State-Size · ICML 2025
Hungry Hungry Hippos: Towards Language Modeling with State Space Models · ICLR 2023
Simple Hardware-Efficient Long Convolutions for Sequence Modeling · ICML 2023
Machine learning › Efficient and distributed learning › automated machine learning
architecture optimization
0.912025
STAR: Synthesis of Tailored Architectures · ICLR 2025
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search
0.912025
STAR: Synthesis of Tailored Architectures · ICLR 2025
Machine learning › Deep learning architectures and training › scaling laws
compute-optimal scaling
0.812024
Mechanistic Design and Scaling of Hybrid Architectures · ICML 2024
Machine learning › Deep learning architectures and training
scaling laws
0.812024
Mechanistic Design and Scaling of Hybrid Architectures · ICML 2024
Machine learning › Efficient and distributed learning › model compression
efficient architecture design
0.712023
Monarch Mixer: A Simple Sub-Quadratic GEMM-Based Architecture · NeurIPS 2023
Machine learning › Efficient and distributed learning › model compression
structured matrices
0.712023
Monarch Mixer: A Simple Sub-Quadratic GEMM-Based Architecture · NeurIPS 2023
Bioinformatics and computational biology › genomics › machine learning for genomics
genomic foundation models
0.712023
HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide Resolution · NeurIPS 2023
Bioinformatics and computational biology › sequence analysis › sequence modeling
genomic sequence modeling
0.712023
HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide Resolution · NeurIPS 2023
Natural language and speech › Language models and text generation › language modeling › long-context language modeling
context utilization
0.312025
Quantifying Memory Utilization with Effective State-Size · ICML 2025
Natural language and speech › Language models and text generation
in-context learning
0.212023
HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide Resolution · NeurIPS 2023
Medical and health informatics › neuroimaging
neuroimaging analysis
0.212022
Self-Supervised Learning of Brain Dynamics from Broad Neuroimaging Data · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

signal processing · 0.9model distillation · 0.9gradient-free optimization · 0.9evolutionary algorithm · 0.9control theory · 0.9synthetic capability unit tests · 0.8scaling law analysis · 0.8state space model · 0.7single-nucleotide tokenization · 0.7pre-training · 0.7implicit convolution · 0.7butterfly decomposition · 0.7IO-aware algorithms · 0.7transfer learning · 0.6self-supervised learning · 0.6causal language modeling · 0.6
YearPublicationVenuePosition
2025 STAR: Synthesis of Tailored Architectures
abstract
Iterative improvement of model architectures is fundamental to deep learning: Transformers first enabled scaling, and recent advances in model hybridization have pushed the quality-efficiency frontier. However, optimizing architectures remains challenging and expensive, with a variety of automated or manual approaches that fall short, due to limited progress in the design of search spaces and due to the simplicity of resulting patterns and heuristics. In this work, we propose a new approach for the synthesis of tailored architectures (STAR). Our approach combines a novel search space based on the theory of linear input-varying systems, supporting a hierarchical numerical encoding into architecture genomes. STAR genomes are automatically refined and recombined with gradient-free, evolutionary algorithms to optimize for multiple model quality and efficiency metrics. Using STAR, we optimize large populations of new architectures, leveraging diverse computational units and interconnection patterns, improving over highly-optimized Transformers and striped hybrid models on the frontier of quality, parameter size, and inference cache for autoregressive language modeling.
Armin W. Thomas, Rom N. Parnichkun, Alexander Amini, Stefano Massaroli, Michael Poli
ICLR1
2025 Quantifying Memory Utilization with Effective State-Size
abstract
As the space of causal sequence modeling architectures continues to grow, the need to develop a general framework for their analysis becomes increasingly important. With this aim, we draw insights from classical signal processing and control theory, to develop a quantitative measure of memory utilization: the internal mechanisms through which a model stores past information to produce future outputs. This metric, which we call effective state-size (ESS), is tailored to the fundamental class of systems with input-invariant and input-varying linear operators, encompassing a variety of computational units such as variants of attention, convolutions, and recurrences. Unlike prior work on memory utilization, which either relies on raw operator visualizations (e.g. attention maps), or simply the total memory capacity (i.e. cache size) of a model, our metrics provide highly interpretable and actionable measurements. In particular, we show how ESS can be leveraged to improve initialization strategies, inform novel regularizers and advance the performance-efficiency frontier through model distillation. Furthermore, we demonstrate that the effect of context delimiters (such as end-of-speech tokens) on ESS highlights cross-architectural differences in how large language models utilize their available memory to recall information. Overall, we find that ESS provides valuable insights into the dynamics that dictate memory utilization, enabling the design of more efficient and effective sequence models.
Rom N. Parnichkun, Neehal Tumma, Armin W. Thomas, Alessandro Moro, Qi An 0001, Taiji Suzuki, Atsushi Yamashita, Michael Poli, Stefano Massaroli
ICML3
2024 Mechanistic Design and Scaling of Hybrid Architectures
abstract
The development of deep learning architectures is a resource-demanding process, due to a vast design space, long prototyping times, and high compute costs associated with at-scale model training and evaluation. We set out to simplify this process by grounding it in an end-to-end mechanistic architecture design (MAD) pipeline, encompassing small-scale capability unit tests predictive of scaling laws. Through a suite of synthetic token manipulation tasks such as compression and recall, designed to probe capabilities, we identify and test new hybrid architectures constructed from a variety of computational primitives. We experimentally validate the resulting architectures via an extensive compute-optimal and a new state-optimal scaling law analysis, training over 500 language models between 70M to 7B parameters. Surprisingly, we find MAD synthetics to correlate with compute-optimal perplexity, enabling accurate evaluation of new architectures via isolated proxy tasks. The new architectures found via MAD, based on simple ideas such as hybridization and sparsity, outperform state-of-the-art Transformer, convolutional, and recurrent architectures (Transformer++, Hyena, Mamba) in scaling, both at compute-optimal budgets and in overtrained regimes. Overall, these results provide evidence that performance on curated synthetic tasks can be predictive of scaling laws, and that an optimal architecture should leverage specialized layers via a hybrid topology.
Michael Poli, Armin W. Thomas, Eric Nguyen, Pragaash Ponnusamy, Björn Deiseroth, Kristian Kersting, Taiji Suzuki, Brian L. Hie, Stefano Ermon, Christopher Ré, Ce Zhang 0001, Stefano Massaroli
ICML2
2023 Hungry Hungry Hippos: Towards Language Modeling with State Space Models
Daniel Y. Fu, Tri Dao, Khaled Saab 0002, Armin W. Thomas, Atri Rudra, Christopher Ré
ICLR4
2023 Simple Hardware-Efficient Long Convolutions for Sequence Modeling
abstract
State space models (SSMs) have high performance on long sequence modeling but require sophisticated initialization techniques and specialized implementations for high quality and runtime performance. We study whether a simple alternative can match SSMs in performance and efficiency: directly learning long convolutions over the sequence. We find that a key requirement to achieving high performance is keeping the convolution kernels smooth. We find that simple interventions-such as squashing the kernel weights-result in smooth kernels and recover SSM performance on a range of tasks including the long range arena, image classification, language modeling, and brain data modeling. Next, we develop FlashButterfly, an IO-aware algorithm to improve the runtime performance of long convolutions. FlashButterfly appeals to classic Butterfly decompositions of the convolution to reduce GPU memory IO and increase FLOP utilization. FlashButterfly speeds up convolutions by 2.2$\times$, and allows us to train on Path256, a challenging task with sequence length 64K, where we set state-of-the-art by 29.1 points while training 7.2$\times$ faster than prior work. Lastly, we introduce an extension to FlashButterfly that learns the coefficients of the Butterfly decomposition, increasing expressivity without increasing runtime. Using this extension, we outperform a Transformer on WikiText103 by 0.2 PPL with 30% fewer parameters.
Daniel Y. Fu, Elliot L. Epstein, Eric Nguyen, Armin W. Thomas, Tri Dao, Atri Rudra, Christopher Ré
ICML4
2023 Monarch Mixer: A Simple Sub-Quadratic GEMM-Based Architecture
abstract
Machine learning models are increasingly being scaled in both sequence length and model dimension to reach longer contexts and better performance. However, existing architectures such as Transformers scale quadratically along both these axes. We ask: are there performant architectures that can scale sub-quadratically along sequence length and model dimension? We introduce Monarch Mixer (M2), a new architecture that uses the same sub-quadratic primitive along both sequence length and model dimension: Monarch matrices, a simple class of expressive structured matrices that captures many linear transforms, achieves high hardware efficiency on GPUs, and scales sub-quadratically. As a proof of concept, we explore the performance of M2 in three domains: non-causal BERT-style language modeling, ViT-style image classification, and causal GPT-style language modeling. For non-causal BERT-style modeling, M2 matches BERT-base and BERT-large in downstream GLUE quality with up to 27% fewer parameters, and achieves up to 9.1$\times$ higher throughput at sequence length 4K. On ImageNet, M2 outperforms ViT-b by 1% in accuracy, with only half the parameters. Causal GPT-style models introduce a technical challenge: enforcing causality via masking introduces a quadratic bottleneck. To alleviate this bottleneck, we develop a novel theoretical view of Monarch matrices based on multivariate polynomial evaluation and interpolation, which lets us parameterize M2 to be causal while remaining sub-quadratic. Using this parameterization, M2 matches GPT-style Transformers at 360M parameters in pretraining perplexity on The PILE—showing for the first time that it may be possible to match Transformer quality without attention or MLPs.
Daniel Y. Fu, Simran Arora, Jessica Grogan, Isys Johnson, Sabri Eyuboglu, Armin W. Thomas, Benjamin Spector, Michael Poli, Atri Rudra, Christopher Ré
NeurIPS6
2023 HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide Resolution
abstract
Genomic (DNA) sequences encode an enormous amount of information for gene regulation and protein synthesis. Similar to natural language models, researchers have proposed foundation models in genomics to learn generalizable features from unlabeled genome data that can then be fine-tuned for downstream tasks such as identifying regulatory elements. Due to the quadratic scaling of attention, previous Transformer-based genomic models have used 512 to 4k tokens as context (<0.001% of the human genome), significantly limiting the modeling of long-range interactions in DNA. In addition, these methods rely on tokenizers or fixed k-mers to aggregate meaningful DNA units, losing single nucleotide resolution (i.e. DNA "characters") where subtle genetic variations can completely alter protein function via single nucleotide polymorphisms (SNPs). Recently, Hyena, a large language model based on implicit convolutions was shown to match attention in quality while allowing longer context lengths and lower time complexity. Leveraging Hyena’s new long-range capabilities, we present HyenaDNA, a genomic foundation model pretrained on the human reference genome with context lengths of up to 1 million tokens at the single nucleotide-level – an up to 500x increase over previous dense attention-based models. HyenaDNA scales sub-quadratically in sequence length (training up to 160x faster than Transformer), uses single nucleotide tokens, and has full global context at each layer. We explore what longer context enables - including the first use of in-context learning in genomics for simple adaptation to novel tasks without updating pretrained model weights. On fine-tuned benchmarks from the Nucleotide Transformer, HyenaDNA reaches state-of-the-art (SotA) on 12 of 18 datasets using a model with orders of magnitude less parameters and pretraining data.1 On the GenomicBenchmarks, HyenaDNA surpasses SotA on 7 of 8 datasets on average by +10 accuracy points. Code at https://github.com/HazyResearch/hyena-dna.
Eric Nguyen, Michael Poli, Marjan Faizi, Armin W. Thomas, Michael Wornow, Callum Birch-Sykes, Stefano Massaroli, Aman Patel, Clayton M. Rabideau, Yoshua Bengio, Stefano Ermon, Christopher Ré, Stephen A. Baccus
NeurIPS4
2022 Self-Supervised Learning of Brain Dynamics from Broad Neuroimaging Data
abstract
Self-supervised learning techniques are celebrating immense success in natural language processing (NLP) by enabling models to learn from broad language data at unprecedented scales. Here, we aim to leverage the success of these techniques for mental state decoding, where researchers aim to identify specific mental states (e.g., the experience of anger or joy) from brain activity. To this end, we devise a set of novel self-supervised learning frameworks for neuroimaging data inspired by prominent learning frameworks in NLP. At their core, these frameworks learn the dynamics of brain activity by modeling sequences of activity akin to how sequences of text are modeled in NLP. We evaluate the frameworks by pre-training models on a broad neuroimaging dataset spanning functional Magnetic Resonance Imaging data from 11,980 experimental runs of 1,726 individuals across 34 datasets, and subsequently adapting the pre-trained models to benchmark mental state decoding datasets. The pre-trained models transfer well, generally outperforming baseline models trained from scratch, while models trained in a learning framework based on causal language modeling clearly outperform the others.
Armin W. Thomas, Christopher Ré, Russell A. Poldrack
NeurIPS1
2022 Gaze-dependent evidence accumulation predicts multi-alternative risky choice behaviour
abstract
Choices are influenced by gaze allocation during deliberation, so that fixating an alternative longer leads to increased probability of choosing it. Gaze-dependent evidence accumulation provides a parsimonious account of choices, response times and gaze-behaviour in many simple decision scenarios. Here, we test whether this framework can also predict more complex context-dependent patterns of choice in a three-alternative risky choice task, where choices and eye movements were subject to attraction and compromise effects. Choices were best described by a gaze-dependent evidence accumulation model, where subjective values of alternatives are discounted while not fixated. Finally, we performed a systematic search over a large model space, allowing us to evaluate the relative contribution of different forms of gaze-dependence and additional mechanisms previously not considered by gaze-dependent accumulation models. Gaze-dependence remained the most important mechanism, but participants with strong attraction effects employed an additional similarity-dependent inhibition mechanism found in other models of multi-alternative multi-attribute choice.
Felix Molter, Armin W. Thomas, Scott A. Huettel, Hauke R. Heekeren, Peter N. C. Mohr
PLoS Comput. Biol.2