VLDB 2026 Research / reviewers in the wild / expert
Jusen Du
dblp:400/4863
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Deep learning architectures and training · 66% Language models and text generation · 34% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
recurrent neural network |
1.7 | 2 | 2025 | Improving Bilinear RNN with Closed-loop Control · NeurIPS 2025 Liger: Linearizing Large Language Models to Gated Recurrent Structures · ICML 2025 |
Machine learning › Deep learning architectures and training
attention mechanism |
1.0 | 1 | 2026 | Native Hybrid Attention for Efficient Sequence Modeling · ACL (1) 2026 |
Machine learning › Deep learning architectures and training › attention mechanism
hybrid attention |
1.0 | 1 | 2026 | Native Hybrid Attention for Efficient Sequence Modeling · ACL (1) 2026 |
Natural language and speech › Language models and text generation
large language model |
0.9 | 1 | 2025 | Liger: Linearizing Large Language Models to Gated Recurrent Structures · ICML 2025 |
Natural language and speech › Language models and text generation › text generation › surface realization
linearization |
0.9 | 1 | 2025 | Liger: Linearizing Large Language Models to Gated Recurrent Structures · ICML 2025 |
Machine learning › Deep learning architectures and training › sequence modeling
efficient sequence modeling |
0.3 | 1 | 2026 | Native Hybrid Attention for Efficient Sequence Modeling · ACL (1) 2026 |
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization
long-context modeling |
0.3 | 1 | 2026 | Native Hybrid Attention for Efficient Sequence Modeling · ACL (1) 2026 |
Mathematical optimization
control theory |
0.3 | 1 | 2025 | Improving Bilinear RNN with Closed-loop Control · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
state feedback · 1.7output feedback · 1.7delta learning rule · 1.7chunk-wise parallel kernel · 1.7softmax attention · 1.0sliding window attention · 1.0linear attention · 1.0low-rank adaptation · 0.9hybrid attention · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Native Hybrid Attention for Efficient Sequence ModelingabstractTransformers excel at sequence modeling but face quadratic complexity, while linear attention offers improved efficiency but often compromises recall accuracy over long contexts.In this work, we introduce Native Hybrid Attention (NHA), a novel hybrid architecture of linear and full attention that integrates both intra & inter-layer hybridization into a unified layer design.NHA maintains longterm context in key-value slots updated by a linear RNN, and augments them with shortterm tokens from a sliding window.A single softmax attention operation is then applied over all keys and values, enabling pertoken and per-head context-dependent weighting without requiring additional fusion parameters.The inter-layer behavior is controlled through a single hyperparameter, the sliding window size, which allows smooth adjustment between purely linear and full attention while keeping all layers structurally uniform.Experimental results show that NHA surpasses Transformers and other hybrid baselines on recall-intensive and commonsense reasoning tasks.Furthermore, pretrained LLMs can be structurally hybridized with NHA, achieving competitive accuracy while delivering significant efficiency gains.Code is available at https://github.com/JusenD/NHA. Jusen Du, Jiaxi Hu, Zhang Tao, Weigao Sun, Yu Cheng 0001 |
ACL (1) | 1 |
| 2025 | Liger: Linearizing Large Language Models to Gated Recurrent StructuresabstractTransformers with linear recurrent modeling offer linear-time training and constant-memory inference. Despite their demonstrated efficiency and performance, pretraining such non-standard architectures from scratch remains costly and risky. The linearization of large language models (LLMs) transforms pretrained standard models into linear recurrent structures, enabling more efficient deployment. However, current linearization methods typically introduce additional feature map modules that require extensive fine-tuning and overlook the gating mechanisms used in state-of-the-art linear recurrent models. To address these issues, this paper presents Liger, short for Linearizing LLMs to gated recurrent structures. Liger is a novel approach for converting pretrained LLMs into gated linear recurrent models without adding extra parameters. It repurposes the pretrained key matrix weights to construct diverse gating mechanisms, facilitating the formation of various gated recurrent structures while avoiding the need to train additional components from scratch. Using lightweight fine-tuning with Low-Rank Adaptation (LoRA), Liger restores the performance of the linearized gated recurrent models to match that of the original LLMs. Additionally, we introduce Liger Attention, an intra-layer hybrid attention mechanism, which significantly recovers 93% of the Transformer-based LLM performance at 0.02% pre-training tokens during the linearization process, achieving competitive results across multiple benchmarks, as validated on models ranging from 1B to 8B parameters. Disen Lan, Weigao Sun, Jiaxi Hu, Jusen Du, Yu Cheng 0001 |
ICML | 4 |
| 2025 | Improving Bilinear RNN with Closed-loop ControlabstractRecent efficient sequence modeling methods, such as Gated DeltaNet, TTT, and RWKV-7, have achieved performance improvements by supervising the recurrent memory management through the Delta learning rule. Unlike previous state-space models (e.g., Mamba) and gated linear attentions (e.g., GLA), these models introduce interactions between the recurrent state and the key vector, resulting in a bilinear recursive structure. In this paper, we first introduce the concept of Bilinear RNNs with a comprehensive analysis on the advantages and limitations of these models. Then based on the closed-loop control theory, we propose a novel Bilinear RNN variant named Comba, which adopts a scalar-plus-low-rank state transition, with both state feedback and output feedback corrections. We also implement a hardware-efficient chunk-wise parallel kernel in Triton and train models with 340M/1.3B parameters on a large-scale corpus. Comba demonstrates its superior performance and computation efficiency on both language modeling and vision tasks. Jiaxi Hu, Yongqi Pan, Jusen Du, Disen Lan, Xiaqiang Tang, Qingsong Wen, Yuxuan Liang 0002, Weigao Sun |
NeurIPS | 3 |