Annan Yu

dblp:302/3962 · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
6since 2021 · last 2025
0009-0007-4827-6830ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Deep learning architectures and training · 100%
Computer graphics and multimedia
1 paper
Image and video processing · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
state space model
3.442025
Block-Biased Mamba for Long-Range Sequence Processing · NeurIPS 2025
HOPE for a Robust Parameterization of Long-memory State Space Models · ICLR 2025
Tuning Frequency Bias of State Space Models · ICLR 2025
Machine learning › Deep learning architectures and training › weight initialization
parameterization and initialization
1.622025
HOPE for a Robust Parameterization of Long-memory State Space Models · ICLR 2025
Robustifying State-space Models for Long Sequences via Approximate Diagonalization · ICLR 2024
Machine learning › Deep learning architectures and training
sequence modeling
1.622025
HOPE for a Robust Parameterization of Long-memory State Space Models · ICLR 2025
Robustifying State-space Models for Long Sequences via Approximate Diagonalization · ICLR 2024
Machine learning › Deep learning architectures and training › sequence modeling
long sequence modeling
1.642025
Tuning Frequency Bias of State Space Models · ICLR 2025
Block-Biased Mamba for Long-Range Sequence Processing · NeurIPS 2025
HOPE for a Robust Parameterization of Long-memory State Space Models · ICLR 2025
Machine learning › Deep learning architectures and training › state space model
mamba
0.912025
Block-Biased Mamba for Long-Range Sequence Processing · NeurIPS 2025
Machine learning › Deep learning architectures and training › training dynamics
frequency bias
0.712023
Tuning Frequency Bias in Neural Network Training with Nonuniform Data · ICLR 2023
Machine learning › Deep learning architectures and training
training dynamics
0.712023
Tuning Frequency Bias in Neural Network Training with Nonuniform Data · ICLR 2023
Image and video processing › image restoration
image denoising
0.312025
Tuning Frequency Bias of State Space Models · ICLR 2025

Methods — techniques the papers use, named apart from their topics

sobolev-norm-based filter · 1.7initialization scaling · 1.7selective dynamics · 0.9nonuniform sampling of transfer functions · 0.9inductive bias analysis · 0.9hankel operator theory · 0.9channel-specific bias · 0.9pseudospectral theory · 0.8perturb-then-diagonalize · 0.8
YearPublicationVenuePosition
2025 Tuning Frequency Bias of State Space Models
abstract
State space models (SSMs) leverage linear, time-invariant (LTI) systems to effectively learn sequences with long-range dependencies. By analyzing the transfer functions of LTI systems, we find that SSMs exhibit an implicit bias toward capturing low-frequency components more effectively than high-frequency ones. This behavior aligns with the broader notion of frequency bias in deep learning model training. We show that the initialization of an SSM assigns it an innate frequency bias and that training the model in a conventional way does not alter this bias. Based on our theory, we propose two mechanisms to tune frequency bias: either by scaling the initialization to tune the inborn frequency bias; or by applying a Sobolev-norm-based filter to adjust the sensitivity of the gradients to high-frequency inputs, which allows us to change the frequency bias via training. Using an image-denoising task, we empirically show that we can strengthen, weaken, or even reverse the frequency bias using both mechanisms. By tuning the frequency bias, we can also improve SSMs' performance on learning long-range sequences, averaging an $88.26\\%$ accuracy on the Long-Range Arena (LRA) benchmark tasks.
Annan Yu, Dongwei Lyu, Soon Hoe Lim, Michael W. Mahoney, N. Benjamin Erichson
ICLR1
2025 HOPE for a Robust Parameterization of Long-memory State Space Models
abstract
State-space models (SSMs) that utilize linear, time-invariant (LTI) systems are known for their effectiveness in learning long sequences. To achieve state-of-the-art performance, an SSM often needs a specifically designed initialization, and the training of state matrices is on a logarithmic scale with a very small learning rate. To understand these choices from a unified perspective, we view SSMs through the lens of Hankel operator theory. Building upon it, we develop a new parameterization scheme, called HOPE, for LTI systems that utilizes Markov parameters within Hankel operators. Our approach helps improve the initialization and training stability, leading to a more robust parameterization. We efficiently implement these innovations by nonuniformly sampling the transfer functions of LTI systems, and they require fewer parameters compared to canonical SSMs. When benchmarked against HiPPO-initialized models such as S4 and S4D, an SSM parameterized by Hankel operators demonstrates improved performance on Long-Range Arena (LRA) tasks. Moreover, our new parameterization endows the SSM with non-decaying memory within a fixed time window, which is empirically corroborated by a sequential CIFAR-10 task with padded noise.
Annan Yu, Michael W. Mahoney, N. Benjamin Erichson
ICLR1
2025 Block-Biased Mamba for Long-Range Sequence Processing
abstract
Mamba extends earlier state space models (SSMs) by introducing input-dependent dynamics, and has demonstrated strong empirical performance across a range of domains, including language modeling, computer vision, and foundation models. However, a surprising weakness remains: despite being built on architectures designed for long-range dependencies, Mamba performs poorly on long-range sequential tasks. Understanding and addressing this gap is important for improving Mamba's universality and versatility. In this work, we analyze Mamba’s limitations through three perspectives: expressiveness, inductive bias, and training stability. Our theoretical results show how Mamba falls short in each of these aspects compared to earlier SSMs such as S4D. To address these issues, we propose $\text{B}\_{2}\text{S}\_{6}$, a simple extension of Mamba's S6 unit that combines block-wise selective dynamics with a channel-specific bias. We prove that these changes equip the model with a better-suited inductive bias and improve its expressiveness and stability. Empirically, $\text{B}\_{2}\text{S}\_{6}$ outperforms S4 and S4D on Long-Range Arena (LRA) tasks while maintaining Mamba's performance on language modeling benchmarks.
Annan Yu, N. Benjamin Erichson
NeurIPS1
2024 Robustifying State-space Models for Long Sequences via Approximate Diagonalization
abstract
State-space models (SSMs) have recently emerged as a framework for learning long-range sequence tasks. An example is the structured state-space sequence (S4) layer, which uses the diagonal-plus-low-rank structure of the HiPPO initialization framework. However, the complicated structure of the S4 layer poses challenges; and, in an effort to address these challenges, models such as S4D and S5 have considered a purely diagonal structure. This choice simplifies the implementation, improves computational efficiency, and allows channel communication. However, diagonalizing the HiPPO framework is itself an ill-posed problem. In this paper, we propose a general solution for this and related ill-posed diagonalization problems in machine learning. We introduce a generic, backward-stable ``perturb-then-diagonalize'' (PTD) methodology, which is based on the pseudospectral theory of non-normal operators, and which may be interpreted as the approximate diagonalization of the non-normal matrices defining SSMs. Based on this, we introduce the S4-PTD and S5-PTD models. Through theoretical analysis of the transfer functions of different initialization schemes, we demonstrate that the S4-PTD/S5-PTD initialization strongly converges to the HiPPO framework, while the S4D/S5 initialization only achieves weak convergences. As a result, our new models show resilience to Fourier-mode noise-perturbed inputs, a crucial property not achieved by the S4D/S5 models. In addition to improved robustness, our S5-PTD model averages 87.6% accuracy on the Long-Range Arena benchmark, demonstrating that the PTD methodology helps to improve the accuracy of deep learning models.
Annan Yu, Arnur Nigmetov, Dmitriy Morozov, Michael W. Mahoney, N. Benjamin Erichson
ICLR1
2023 Tuning Frequency Bias in Neural Network Training with Nonuniform Data
Annan Yu, Yunan Yang, Alex Townsend
ICLR1
2022 Approximation by polynomial splines on curved triangulations
Larry L. Schumaker, Annan Yu
Comput. Aided Geom. Des.2