Disen Lan

dblp:371/6233 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0004-5626-7641ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Deep learning architectures and training · 43% Graph learning · 23% Language models and text generation · 20%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
recurrent neural network
1.722025
Improving Bilinear RNN with Closed-loop Control · NeurIPS 2025
Liger: Linearizing Large Language Models to Gated Recurrent Structures · ICML 2025
Machine learning › Time series and sequential data › time series analysis
time series forecasting
1.222026
TimeFilter: Patch-Specific Spatial-Temporal Graph Filtration for Time Series Forecasting · ICML 2025
How to Train Your Mamba for Time Series Forecasting · KDD (1) 2026
Machine learning › Deep learning architectures and training
sequence modeling
1.012026
How to Train Your Mamba for Time Series Forecasting · KDD (1) 2026
Machine learning › Deep learning architectures and training
state space model
1.012026
How to Train Your Mamba for Time Series Forecasting · KDD (1) 2026
Machine learning › Graph learning
graph neural network
0.912025
TimeFilter: Patch-Specific Spatial-Temporal Graph Filtration for Time Series Forecasting · ICML 2025
Natural language and speech › Language models and text generation
large language model
0.912025
Liger: Linearizing Large Language Models to Gated Recurrent Structures · ICML 2025
Natural language and speech › Language models and text generation › text generation › surface realization
linearization
0.912025
Liger: Linearizing Large Language Models to Gated Recurrent Structures · ICML 2025
Machine learning › Graph learning › graph neural network
spatio-temporal graph
0.912025
TimeFilter: Patch-Specific Spatial-Temporal Graph Filtration for Time Series Forecasting · ICML 2025
Data mining › time series analysis › time series classification
shapelet learning
0.812024
Diffusion Language-Shapelets for Semi-supervised Time-Series Classification · AAAI 2024
Data mining › time series analysis
time series classification
0.812024
Diffusion Language-Shapelets for Semi-supervised Time-Series Classification · AAAI 2024
Machine learning › Graph learning › graph signal processing
graph filter
0.312025
TimeFilter: Patch-Specific Spatial-Temporal Graph Filtration for Time Series Forecasting · ICML 2025
Mathematical optimization
control theory
0.312025
Improving Bilinear RNN with Closed-loop Control · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

output feedback · 1.7delta learning rule · 1.7chunk-wise parallel kernel · 1.7spectral analysis · 1.0ablation study · 1.0state feedback · 0.9patch-specific filtering · 0.9low-rank adaptation · 0.9hybrid attention · 0.9graph neural network · 0.9channel clustering · 0.9natural language description · 0.8diffusion model · 0.8contrastive learning · 0.8
YearPublicationVenuePosition
2026 How to Train Your Mamba for Time Series Forecasting
abstract
State Space Models (SSMs) have emerged as a powerful framework for sequence modeling in recent years. By approximating continuous dynamical systems and applying discretization techniques, SSMs are particularly well-suited for modeling time-series data. However, despite their growing popularity, most existing applications of SSMs in time-series forecasting treat the models as black boxes. Besides, the underlying mechanisms that contribute to their effectiveness remain unclear, and common claims regarding their advantages in efficiency and expressiveness are not fully substantiated. To address these gaps, this paper establishes a theoretical connection between SSMs and classical spectral transformations from signal processing, thereby providing a more interpretable foundation. Furthermore, we conduct comprehensive ablation studies to examine the properties of different SSM configurations. Our goal is to offer both theoretical insight and empirical guidance for future research on SSM-based approaches in time-series forecasting.
Jiaxi Hu, Disen Lan, Ziyu Zhou 0003, Gefeng Luo, Qingsong Wen, Yuxuan Liang 0002
KDD (1)2
2025 TimeFilter: Patch-Specific Spatial-Temporal Graph Filtration for Time Series Forecasting
abstract
Time series forecasting methods generally fall into two main categories: Channel Independent (CI) and Channel Dependent (CD) strategies. While CI overlooks important covariate relationships, CD captures all dependencies without distinction, introducing noise and reducing generalization. Recent advances in Channel Clustering (CC) aim to refine dependency modeling by grouping channels with similar characteristics and applying tailored modeling techniques. However, coarse-grained clustering struggles to capture complex, time-varying interactions effectively. To address these challenges, we propose TimeFilter, a GNN-based framework for adaptive and fine-grained dependency modeling. After constructing the graph from the input sequence, TimeFilter refines the learned spatial-temporal dependencies by filtering out irrelevant correlations while preserving the most critical ones in a patch-specific manner. Extensive experiments on 13 real-world datasets from diverse application domains demonstrate the state-of-the-art performance of TimeFilter. The code is available at https://github.com/TROUBADOUR000/TimeFilter.
Yifan Hu 0006, Guibin Zhang, Peiyuan Liu, Disen Lan, Naiqi Li, Dawei Cheng, Tao Dai 0001, Shutao Xia, Shirui Pan
ICML4
2025 Liger: Linearizing Large Language Models to Gated Recurrent Structures
abstract
Transformers with linear recurrent modeling offer linear-time training and constant-memory inference. Despite their demonstrated efficiency and performance, pretraining such non-standard architectures from scratch remains costly and risky. The linearization of large language models (LLMs) transforms pretrained standard models into linear recurrent structures, enabling more efficient deployment. However, current linearization methods typically introduce additional feature map modules that require extensive fine-tuning and overlook the gating mechanisms used in state-of-the-art linear recurrent models. To address these issues, this paper presents Liger, short for Linearizing LLMs to gated recurrent structures. Liger is a novel approach for converting pretrained LLMs into gated linear recurrent models without adding extra parameters. It repurposes the pretrained key matrix weights to construct diverse gating mechanisms, facilitating the formation of various gated recurrent structures while avoiding the need to train additional components from scratch. Using lightweight fine-tuning with Low-Rank Adaptation (LoRA), Liger restores the performance of the linearized gated recurrent models to match that of the original LLMs. Additionally, we introduce Liger Attention, an intra-layer hybrid attention mechanism, which significantly recovers 93% of the Transformer-based LLM performance at 0.02% pre-training tokens during the linearization process, achieving competitive results across multiple benchmarks, as validated on models ranging from 1B to 8B parameters.
Disen Lan, Weigao Sun, Jiaxi Hu, Jusen Du, Yu Cheng 0001
ICML1
2025 Improving Bilinear RNN with Closed-loop Control
abstract
Recent efficient sequence modeling methods, such as Gated DeltaNet, TTT, and RWKV-7, have achieved performance improvements by supervising the recurrent memory management through the Delta learning rule. Unlike previous state-space models (e.g., Mamba) and gated linear attentions (e.g., GLA), these models introduce interactions between the recurrent state and the key vector, resulting in a bilinear recursive structure. In this paper, we first introduce the concept of Bilinear RNNs with a comprehensive analysis on the advantages and limitations of these models. Then based on the closed-loop control theory, we propose a novel Bilinear RNN variant named Comba, which adopts a scalar-plus-low-rank state transition, with both state feedback and output feedback corrections. We also implement a hardware-efficient chunk-wise parallel kernel in Triton and train models with 340M/1.3B parameters on a large-scale corpus. Comba demonstrates its superior performance and computation efficiency on both language modeling and vision tasks.
Jiaxi Hu, Yongqi Pan, Jusen Du, Disen Lan, Xiaqiang Tang, Qingsong Wen, Yuxuan Liang 0002, Weigao Sun
NeurIPS4
2024 Diffusion Language-Shapelets for Semi-supervised Time-Series Classification
abstract
Semi-supervised time-series classification could effectively alleviate the issue of lacking labeled data. However, existing approaches usually ignore model interpretability, making it difficult for humans to understand the principles behind the predictions of a model. Shapelets are a set of discriminative subsequences that show high interpretability in time series classification tasks. Shapelet learning-based methods have demonstrated promising classification performance. Unfortunately, without enough labeled data, the shapelets learned by existing methods are often poorly discriminative, and even dissimilar to any subsequence of the original time series. To address this issue, we propose the Diffusion Language-Shapelets model (DiffShape) for semi-supervised time series classification. In DiffShape, a self-supervised diffusion learning mechanism is designed, which uses real subsequences as a condition. This helps to increase the similarity between the learned shapelets and real subsequences by using a large amount of unlabeled data. Furthermore, we introduce a contrastive language-shapelets learning strategy that improves the discriminability of the learned shapelets by incorporating the natural language descriptions of the time series. Experiments have been conducted on the UCR time series archive, and the results reveal that the proposed DiffShape method achieves state-of-the-art performance and exhibits superior interpretability over baselines.
Zhen Liu 0023, Wenbin Pei, Disen Lan, Qianli Ma 0001
AAAI3