VLDB 2026 Research / reviewers in the wild / expert
Rom N. Parnichkun
dblp:359/5796
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Deep learning architectures and training · 56% Efficient and distributed learning · 38% Language models and text generation · 6% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
sequence modeling |
2.3 | 3 | 2025 | Quantifying Memory Utilization with Effective State-Size · ICML 2025 State-Free Inference of State-Space Models: The *Transfer Function* Approach · ICML 2024 Laughing Hyena Distillery: Extracting Compact Recurrences From Convolutions · NeurIPS 2023 |
Machine learning › Deep learning architectures and training
state space model |
2.3 | 3 | 2025 | Quantifying Memory Utilization with Effective State-Size · ICML 2025 State-Free Inference of State-Space Models: The *Transfer Function* Approach · ICML 2024 Laughing Hyena Distillery: Extracting Compact Recurrences From Convolutions · NeurIPS 2023 |
Machine learning › Efficient and distributed learning › automated machine learning
architecture optimization |
0.9 | 1 | 2025 | STAR: Synthesis of Tailored Architectures · ICLR 2025 |
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search |
0.9 | 1 | 2025 | STAR: Synthesis of Tailored Architectures · ICLR 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.7 | 1 | 2023 | Laughing Hyena Distillery: Extracting Compact Recurrences From Convolutions · NeurIPS 2023 |
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
model distillation |
0.7 | 1 | 2023 | Laughing Hyena Distillery: Extracting Compact Recurrences From Convolutions · NeurIPS 2023 |
Natural language and speech › Language models and text generation › language modeling › long-context language modeling
context utilization |
0.3 | 1 | 2025 | Quantifying Memory Utilization with Effective State-Size · ICML 2025 |
Natural language and speech › Language models and text generation
language modeling |
0.2 | 1 | 2024 | State-Free Inference of State-Space Models: The *Transfer Function* Approach · ICML 2024 |
Methods — techniques the papers use, named apart from their topics
signal processing · 0.9model distillation · 0.9gradient-free optimization · 0.9evolutionary algorithm · 0.9control theory · 0.9transfer function · 0.8sequence parallel inference · 0.8fast fourier transform · 0.8rational interpolation · 0.7model order reduction · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | STAR: Synthesis of Tailored ArchitecturesabstractIterative improvement of model architectures is fundamental to deep learning: Transformers first enabled scaling, and recent advances in model hybridization have pushed the quality-efficiency frontier. However, optimizing architectures remains challenging and expensive, with a variety of automated or manual approaches that fall short, due to limited progress in the design of search spaces and due to the simplicity of resulting patterns and heuristics. In this work, we propose a new approach for the synthesis of tailored architectures (STAR). Our approach combines a novel search space based on the theory of linear input-varying systems, supporting a hierarchical numerical encoding into architecture genomes. STAR genomes are automatically refined and recombined with gradient-free, evolutionary algorithms to optimize for multiple model quality and efficiency metrics. Using STAR, we optimize large populations of new architectures, leveraging diverse computational units and interconnection patterns, improving over highly-optimized Transformers and striped hybrid models on the frontier of quality, parameter size, and inference cache for autoregressive language modeling. Armin W. Thomas, Rom N. Parnichkun, Alexander Amini, Stefano Massaroli, Michael Poli |
ICLR | 2 |
| 2025 | Quantifying Memory Utilization with Effective State-SizeabstractAs the space of causal sequence modeling architectures continues to grow, the need to develop a general framework for their analysis becomes increasingly important. With this aim, we draw insights from classical signal processing and control theory, to develop a quantitative measure of memory utilization: the internal mechanisms through which a model stores past information to produce future outputs. This metric, which we call effective state-size (ESS), is tailored to the fundamental class of systems with input-invariant and input-varying linear operators, encompassing a variety of computational units such as variants of attention, convolutions, and recurrences. Unlike prior work on memory utilization, which either relies on raw operator visualizations (e.g. attention maps), or simply the total memory capacity (i.e. cache size) of a model, our metrics provide highly interpretable and actionable measurements. In particular, we show how ESS can be leveraged to improve initialization strategies, inform novel regularizers and advance the performance-efficiency frontier through model distillation. Furthermore, we demonstrate that the effect of context delimiters (such as end-of-speech tokens) on ESS highlights cross-architectural differences in how large language models utilize their available memory to recall information. Overall, we find that ESS provides valuable insights into the dynamics that dictate memory utilization, enabling the design of more efficient and effective sequence models. Rom N. Parnichkun, Neehal Tumma, Armin W. Thomas, Alessandro Moro, Qi An 0001, Taiji Suzuki, Atsushi Yamashita, Michael Poli, Stefano Massaroli |
ICML | 1 |
| 2024 | State-Free Inference of State-Space Models: The *Transfer Function* ApproachabstractWe approach designing a state-space model for deep learning applications through its dual representation, the transfer function, and uncover a highly efficient sequence parallel inference algorithm that is state-free: unlike other proposed algorithms, state-free inference does not incur any significant memory or computational cost with an increase in state size. We achieve this using properties of the proposed frequency domain transfer function parametrization, which enables direct computation of its corresponding convolutional kernel’s spectrum via a single Fast Fourier Transform. Our experimental results across multiple sequence lengths and state sizes illustrates, on average, a 35% training speed improvement over S4 layers – parametrized in time-domain – on the Long Range Arena benchmark, while delivering state-of-the-art downstream performances over other attention-free approaches. Moreover, we report improved perplexity in language modeling over a long convolutional Hyena baseline, by simply introducing our transfer function parametrization. Our code is available at https://github.com/ruke1ire/RTF. Rom N. Parnichkun, Stefano Massaroli, Alessandro Moro, Jimmy T. H. Smith, Ramin M. Hasani, Mathias Lechner, Qi An 0001, Christopher Ré, Hajime Asama, Stefano Ermon, Taiji Suzuki, Michael Poli, Atsushi Yamashita |
ICML | 1 |
| 2023 | Laughing Hyena Distillery: Extracting Compact Recurrences From ConvolutionsabstractRecent advances in attention-free sequence models rely on convolutions as alternatives to the attention operator at the core of Transformers. In particular, long convolution sequence models have achieved state-of-the-art performance in many domains, but incur a significant cost during auto-regressive inference workloads -- naively requiring a full pass (or caching of activations) over the input sequence for each generated token -- similarly to attention-based models. In this paper, we seek to enable $\mathcal O(1)$ compute and memory cost per token in any pre-trained long convolution architecture to reduce memory footprint and increase throughput during generation. Concretely, our methods consist in extracting low-dimensional linear state-space models from each convolution layer, building upon rational interpolation and model-order reduction techniques. We further introduce architectural improvements to convolution-based layers such as Hyena: by weight-tying the filters across channels into heads, we achieve higher pre-training quality and reduce the number of filters to be distilled. The resulting model achieves 10x higher throughput than Transformers and 1.5x higher than Hyena at 1.3B parameters, without any loss in quality after distillation. Stefano Massaroli, Michael Poli, Daniel Y. Fu, Hermann Kumbong, Rom N. Parnichkun, David W. Romero, Aman Timalsina, Quinn McIntyre, Beidi Chen, Atri Rudra, Ce Zhang 0001, Christopher Ré, Stefano Ermon, Yoshua Bengio |
NeurIPS | 5 |