EDBT 2026 Demo / reviewers in the wild / expert
Zhen Qin 0003
dblp:06/864-3
· DBLP profile ↗
14ranked-venue papers
7as first author
14since 2021 · last 2026
0009-0005-9856-0843ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 7 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Spatial Decay for Vision TransformersabstractVision Transformers (ViTs) have revolutionized computer vision, yet their self-attention mechanism lacks explicit spatial inductive biases, leading to suboptimal performance on spatially-structured tasks. Existing approaches introduce data-independent spatial decay based on fixed distance metrics, applying uniform attention weighting regardless of image content and limiting adaptability to diverse visual scenarios. Inspired by recent advances in large language models where content-aware gating mechanisms (e.g., GLA, HGRN2, FOX) significantly outperform static alternatives, we present the first successful adaptation of data-dependent spatial decay to 2D vision transformers. We introduce Spatial Decay Transformer (SDT), featuring a novel Context-Aware Gating (CAG) mechanism that generates dynamic, data-dependent decay for patch interactions. Our approach learns to modulate spatial attention based on both content relevance and spatial proximity. We address the fundamental challenge of 1D-to-2D adaptation through a unified spatial-content fusion framework that integrates manhattan distance-based spatial priors with learned content representations. Extensive experiments on ImageNet-1K classification and generation tasks demonstrate consistent improvements over strong baselines. Our work establishes data-dependent spatial decay as a new paradigm for enhancing spatial attention in vision transformers. Yuxin Mao, Zhen Qin 0003, Jinxing Zhou, Bin Fan 0002, Jing Zhang 0052, Yiran Zhong, Yuchao Dai |
AAAI | 2 |
| 2025 | Deep Non-Rigid Structure-from-Motion Revisited: Canonicalization and Sequence ModelingabstractNon-Rigid Structure-from-Motion (NRSfM) is a classic 3D vision problem, where a 2D sequence is taken as input to estimate the corresponding 3D sequence. Recently, the deep neural networks have greatly advanced the task of NRSfM. However, existing deep NRSfM methods still have limitations in handling the inherent sequence property and motion ambiguity associated with the NRSfM problem. In this paper, we revisit deep NRSfM from two perspectives to address the limitations of current deep NRSfM methods : (1) canonicalization and (2) sequence modeling. We propose an easy-to-implement per-sequence canonicalization method as opposed to the previous per-dataset canonicalization approaches. With this in mind, we propose a sequence modeling method that combines temporal information and subspace constraint. As a result, we have achieved a more optimal NRSfM reconstruction pipeline compared to previous efforts. The effectiveness of our method is verified by testing the sequence-to-sequence deep NRSfM pipeline with corresponding regularization modules on several commonly used datasets. Zhen Qin 0003, Yiran Zhong, Yuchao Dai |
AAAI | 3 |
| 2024 | Exploring Transformer ExtrapolationabstractLength extrapolation has attracted considerable attention recently since it allows transformers to be tested on longer sequences than those used in training. Previous research has shown that this property can be attained by using carefully designed Relative Positional Encodings (RPEs). While these methods perform well on a variety of corpora, the conditions for length extrapolation have yet to be investigated. This paper attempts to determine what types of RPEs allow for length extrapolation through a thorough mathematical and empirical analysis. We discover that a transformer is certain to possess this property as long as the series that corresponds to the RPE's exponential converges. Two practices are derived from the conditions and examined in language modeling tasks on a variety of corpora. As a bonus from the conditions, we derive a new Theoretical Receptive Field (TRF) to measure the receptive field of RPEs without taking any training steps. Extensive experiments are conducted on the Wikitext-103, Books, Github, and WikiBook datasets to demonstrate the viability of our discovered conditions. We also compare TRF to Empirical Receptive Field (ERF) across different models, showing consistently matched trends on these datasets. Code is released at: https://github.com/OpenNLPLab/Rpe. Zhen Qin 0003, Yiran Zhong |
AAAI | 1 |
| 2024 | Scaling Laws for Linear Complexity Language ModelsabstractThe interest in linear complexity models for large language models is on the rise, although their scaling capacity remains uncertain.In this study, we present the scaling laws for linear complexity language models to establish a foundation for their scalability.Specifically, we examine the scaling behaviors of three efficient linear architectures.These include TNL (Qin et al., 2024c), a linear attention model with data-independent decay; HGRN2 (Qin et al., 2024e), a linear RNN with data-dependent decay; and cosFormer2 (Qin et al., 2022b(Qin et al., , 2024a)), a linear attention model without decay.We also include LLaMA as a baseline architecture for comparison with softmax attention.These models were trained with six variants, ranging from 70M to 7B parameters on a 300B-token corpus, and evaluated with a total of 1,376 intermediate checkpoints on various downstream tasks.These tasks include validation loss, commonsense reasoning, and information retrieval and generation.The study reveals that existing linear complexity language models exhibit similar scaling capabilities as conventional transformer-based models while also demonstrating superior linguistic proficiency and knowledge retention. Xuyang Shen, Dong Li 0033, Ruitao Leng, Zhen Qin 0003, Weigao Sun, Yiran Zhong |
EMNLP | 4 |
| 2024 | CO2: Efficient Distributed Training with Full Communication-Computation OverlapabstractThe fundamental success of large language models hinges upon the efficacious implementation of large-scale distributed training techniques. Nevertheless, building a vast, high-performance cluster featuring high-speed communication interconnectivity is prohibitively costly, and accessible only to prominent entities. In this work, we aim to lower this barrier and democratize large-scale training with limited bandwidth clusters. We propose a new approach called CO2 that introduces local-updating and asynchronous communication to the distributed data-parallel training, thereby facilitating the full overlap of COmmunication with COmputation. CO2 is able to attain a high scalability even on extensive multi-node clusters constrained by very limited communication bandwidth. We further propose the staleness gap penalty and outer momentum clipping techniques together with CO2 to bolster its convergence and training stability. Besides, CO2 exhibits seamless integration with well-established ZeRO-series optimizers which mitigate memory consumption of model states with large model training. We also provide a mathematical proof of convergence, accompanied by the establishment of a stringent upper bound. Furthermore, we validate our findings through an extensive set of practical experiments encompassing a wide range of tasks in the fields of computer vision and natural language processing. These experiments serve to demonstrate the capabilities of CO2 in terms of convergence, generalization, and scalability when deployed across configurations comprising up to 128 A100 GPUs. The outcomes emphasize the outstanding capacity of CO2 to hugely improve scalability, no matter on clusters with 800Gbps RDMA or 80Gbps TCP/IP inter-node connections. Weigao Sun, Zhen Qin 0003, Weixuan Sun, Shidi Li, Dong Li 0033, Xuyang Shen, Yu Qiao 0001, Yiran Zhong |
ICLR | 2 |
| 2024 | Various Lengths, Constant Speed: Efficient Language Modeling with Lightning AttentionabstractWe present Lightning Attention, the first linear attention implementation that maintains a constant training speed for various sequence lengths under fixed memory consumption. Due to the issue with cumulative summation operations (cumsum), previous linear attention implementations cannot achieve their theoretical advantage in a casual setting. However, this issue can be effectively solved by utilizing different attention calculation strategies to compute the different parts of attention. Specifically, we split the attention calculation into intra-blocks and inter-blocks and use conventional attention computation for intra-blocks and linear attention kernel tricks for inter-blocks. This eliminates the need for cumsum in the linear attention calculation. Furthermore, a tiling technique is adopted through both forward and backward procedures to take full advantage of the GPU hardware. To enhance accuracy while preserving efficacy, we introduce TransNormerLLM (TNL), a new architecture that is tailored to our lightning attention. We conduct rigorous testing on standard and self-collected datasets with varying model sizes and sequence lengths. TNL is notably more efficient than other language models. In addition, benchmark results indicate that TNL performs on par with state-of-the-art LLMs utilizing conventional transformer structures. The source code is released at github.com/OpenNLPLab/TransnormerLLM. Zhen Qin 0003, Weigao Sun, Dong Li 0033, Xuyang Shen, Weixuan Sun, Yiran Zhong |
ICML | 1 |
| 2024 | TAVGBench: Benchmarking Text to Audible-Video GenerationabstractThe Text to Audible-Video Generation (TAVG) task involves generating videos with accompanying audio based on text descriptions. Achieving this requires skillful alignment of both audio and video elements. To support research in this field, we have developed a comprehensive Text to Audible-Video Generation Benchmark (TAVGBench), which contains over 1.7 million clips with a total duration of 11.8 thousand hours. We propose an automatic annotation pipeline to ensure each audible video has detailed descriptions for both its audio and video contents. We also introduce the Audio-Visual Harmoni score (AVHScore) to provide a quantitative measure of the alignment between the generated audio and video modalities. Additionally, we present a baseline model for TAVG called TAVDiffusion, which uses a two-stream latent diffusion model to provide a fundamental starting point for further research in this area. We achieve the alignment of audio and video by employing cross-attention and contrastive learning. Through extensive experiments and evaluations on TAVGBench, we demonstrate the effectiveness of our proposed model under both conventional metrics and our proposed metrics. The dataset and code can be found on this page https://npucvr.github.io/TAVGBench/ and on github https://github.com/OpenNLPLab/TAVGBench. Yuxin Mao, Xuyang Shen, Jing Zhang 0052, Zhen Qin 0003, Jinxing Zhou, Mochu Xiang, Yiran Zhong, Yuchao Dai |
ACM Multimedia | 4 |
| 2023 | Fine-grained Audible Video DescriptionabstractWe explore a new task for audio-visual-language modeling called fine-grained audible video description (FAVD). It aims to provide detailed textual descriptions for the given audible videos, including the appearance and spatial locations of each object, the actions of moving objects, and the sounds in videos. Existing visual-language modeling tasks often concentrate on visual cues in videos while undervaluing the language and audio modalities. On the other hand, FAVD requires not only audio-visual-language modeling skills but also paragraph-level language generation abilities. We construct the first fine-grained audible video description benchmark (FAVDBench) to facilitate this research. For each video clip, we first provide a one-sentence summary of the video, i.e., the caption, followed by 4–6 sentences describing the visual details and 1–2 audio-related descriptions at the end. The descriptions are provided in both English and Chinese. We create two new metrics for this task: an EntityScore to gauge the completeness of entities in the visual descriptions, and an AudioScore to assess the audio descriptions. As a preliminary approach to this task, we propose an audio-visual-language transformer that extends existing video captioning model with an additional audio branch. We combine the masked language modeling and auto-regressive language modeling losses to optimize our model so that it can produce paragraph-level descriptions. We illustrate the efficiency of our model in audio-visual-language modeling by evaluating it against the proposed benchmark using both conventional captioning metrics and our proposed metrics. We further put our benchmark to the test in video generation models, demonstrating that employing fine-grained video descriptions can create more intricate videos than using captions. Code and dataset are available at https://github.com/OpenNLPLab/FAVDBench. Our online benchmark is available at www.avlbench.opennlplab.cn. Xuyang Shen, Dong Li 0033, Jinxing Zhou, Zhen Qin 0003, Xiaodong Han, Aixuan Li, Yuchao Dai, Lingpeng Kong, Meng Wang 0001, Yu Qiao 0001, Yiran Zhong |
CVPR | 4 |
| 2023 | Accelerating Toeplitz Neural Network with Constant-time Inference Complexityabstractgithub.com/OpenNLPLab/ETSC-Exact-Toeplitz-to Zhen Qin 0003, Yiran Zhong |
EMNLP | 1 |
| 2023 | Toeplitz Neural Network for Sequence Modeling
Zhen Qin 0003, Xiaodong Han, Weixuan Sun, Dong Li 0033, Dongxu Li 0003, Yuchao Dai, Lingpeng Kong, Yiran Zhong |
ICLR | 1 |
| 2023 | Hierarchically Gated Recurrent Neural Network for Sequence ModelingabstractTransformers have surpassed RNNs in popularity due to their superior abilities in parallel training and long-term dependency modeling.
Recently, there has been a renewed interest in using linear RNNs for efficient sequence modeling.
These linear RNNs often employ gating mechanisms in the output of the linear recurrence layer while ignoring the significance of using forget gates within the recurrence. In this paper, we propose a gated linear RNN model dubbed Hierarchically Gated Recurrent Neural Network (HGRN), which includes forget gates that are lower bounded by a learnable value. The lower bound increases monotonically when moving up layers. This allows the upper layers to model long-term dependencies and the lower layers to model more local, short-term dependencies. Experiments on language modeling, image classification, and long-range arena benchmarks showcase the efficiency and effectiveness of our proposed model. The source code is available at https://github.com/OpenNLPLab/HGRN. Zhen Qin 0003, Yiran Zhong |
NeurIPS | 1 |
| 2023 | Vicinity Vision TransformerabstractVision transformers have shown great success on numerous computer vision tasks. However, their central component, softmax attention, prohibits vision transformers from scaling up to high-resolution images, due to both the computational complexity and memory footprint being quadratic. Linear attention was introduced in natural language processing (NLP) which reorders the self-attention mechanism to mitigate a similar issue, but directly applying existing linear attention to vision may not lead to satisfactory results. We investigate this problem and point out that existing linear attention methods ignore an inductive bias in vision tasks, i.e., 2D locality. In this article, we propose Vicinity Attention, which is a type of linear attention that integrates 2D locality. Specifically, for each image patch, we adjust its attention weight based on its 2D Manhattan distance from its neighbouring patches. In this case, we achieve 2D locality in a linear complexity where the neighbouring image patches receive stronger attention than far away patches. In addition, we propose a novel Vicinity Attention Block that is comprised of Feature Reduction Attention (FRA) and Feature Preserving Connection (FPC) in order to address the computational bottleneck of linear attention approaches, including our Vicinity Attention, whose complexity grows quadratically with respect to the feature dimension. The Vicinity Attention Block computes attention in a compressed feature space with an extra skip connection to retrieve the original feature distribution. We experimentally validate that the block further reduces computation without degenerating the accuracy. Finally, to validate the proposed methods, we build a linear vision transformer backbone named Vicinity Vision Transformer (VVT). Targeting general vision tasks, we build VVT in a pyramid structure with progressively reduced sequence length. We perform extensive experiments on CIFAR-100, ImageNet-1 k, and ADE20 K datasets to validate the effectiveness of our method. Our method has a slower growth rate in terms of computational overhead than previous transformer-based and convolution-based networks when the input resolution increases. In particular, our approach achieves state-of-the-art image classification accuracy with 50% fewer parameters than previous approaches. Weixuan Sun, Zhen Qin 0003, Yi Zhang 0137, Kaihao Zhang, Nick Barnes, Stanley T. Birchfield, Lingpeng Kong, Yiran Zhong |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | The Devil in Linear TransformerabstractLinear transformers aim to reduce the quadratic space-time complexity of vanilla transformers.However, they usually suffer from degraded performances on various tasks and corpora.In this paper, we examine existing kernel-based linear transformers and identify two key issues that lead to such performance gaps: 1) unbounded gradients in the attention computation adversely impact the convergence of linear transformer models; 2) attention dilution which trivially distributes attention scores over long sequences while neglecting neighbouring structures.To address these issues, we first identify that the scaling of attention matrices is the devil in unbounded gradients, which turns out unnecessary in linear attention as we show theoretically and empirically.To this end, we propose a new linear attention that replaces the scaling operation with a normalization to stabilize gradients.For the issue of attention dilution, we leverage a diagonal attention to confine attention to only neighbouring tokens in early layers.Benefiting from the stable gradients and improved attention, our new linear transformer model, TRANSNORMER, demonstrates superior performance on text classification and language modeling tasks, as well as on the challenging Long-Range Arena benchmark, surpassing vanilla transformer and existing linear variants by a clear margin while being significantly more space-time efficient.The code is available at TRANSNORMER. Zhen Qin 0003, Xiaodong Han, Weixuan Sun, Dongxu Li 0003, Lingpeng Kong, Nick Barnes, Yiran Zhong |
EMNLP | 1 |
| 2022 | cosFormer: Rethinking Softmax In Attention
Zhen Qin 0003, Weixuan Sun, Dongxu Li 0003, Yunshen Wei, Baohong Lv, Lingpeng Kong, Yiran Zhong |
ICLR | 1 |