EDBT 2026 Demo / reviewers in the wild / expert
Dongchen Han
dblp:136/2521
· DBLP profile ↗
16ranked-venue papers
8as first author
15since 2021 · last 2026
0009-0009-3431-6189ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 7 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Vision Transformers Are Circulant Attention LearnersabstractThe self-attention mechanism has been a key factor in the advancement of vision Transformers. However, its quadratic complexity imposes a heavy computational burden in high-resolution scenarios, restricting the practical application. Previous methods attempt to mitigate this issue by introducing handcrafted patterns such as locality or sparsity, which inevitably compromise model capacity. In this paper, we present a novel attention paradigm termed Circulant Attention by exploiting the inherent efficient pattern of self-attention. Specifically, we first identify that the self-attention matrix in vision Transformers often approximates the Block Circulant matrix with Circulant Blocks (BCCB), a kind of structured matrix whose multiplication with other matrices can be performed in O(NlogN) time. Leveraging this interesting pattern, we explicitly model the attention map as its nearest BCCB matrix and propose an efficient computation algorithm for fast calculation. The resulting approach closely mirrors vanilla self-attention, differing only in its use of BCCB matrices. Since our design is inspired by the inherent efficient paradigm, it not only delivers O(NlogN) computation complexity, but also largely maintains the capacity of standard self-attention. Extensive experiments on diverse visual tasks demonstrate the effectiveness of our approach, establishing circulant attention as a promising alternative to self-attention for vision Transformer architectures. Dongchen Han, Gao Huang 0001 |
AAAI | 1 |
| 2026 | Boosting adversarial transferability of vision-language pre-trained models via optimal transport
Simeng Qin, Sensen Gao, Dongchen Han, Xiaojun Jia, Yang Bai 0011, Jindong Gu, Xiaochun Cao |
Pattern Recognit. | 4 |
| 2025 | Linear Differential Vision Transformer: Learning Visual Contrasts via Pairwise DifferentialsabstractVision Transformers (ViTs) have become a universal backbone for both image recognition and image generation. Yet their Multi–Head Self–Attention (MHSA) layer still performs a quadratic query–key interaction for \emph{every} token pair, spending the bulk of computation on visually weak or redundant correlations. We introduce \emph{Visual–Contrast Attention} (VCA), a drop-in replacement for MHSA that injects an explicit notion of discrimination while reducing the theoretical complexity from $\mathcal{O}(N^{2}C)$ to $\mathcal{O}(N n C)$ with $n\!\ll\!N$. VCA first distils each head’s dense query field into a handful of spatially pooled \emph{visual–contrast tokens}, then splits them into a learnable \emph{positive} and \emph{negative} stream whose differential interaction highlights what truly separates one region from another. The module adds fewer than $0.3$\,M parameters to a DeiT-Tiny backbone, requires no extra FLOPs, and is wholly architecture-agnostic. Empirically, VCA lifts DeiT-Tiny top-1 accuracy on ImageNet-1K from $72.2\%$ to \textbf{$75.6\%$} (+$3.4$) and improves three strong hierarchical ViTs by up to $3.1$\%, while in class-conditional ImageNet generation it lowers FID-50K by $2.1$ to $5.2$ points across both diffusion (DiT) and flow (SiT) models. Extensive ablations confirm that (i) spatial pooling supplies low-variance global cues, (ii) dual positional embeddings are indispensable for contrastive reasoning, and (iii) combining the two in both stages yields the strongest synergy. VCA therefore offers a simple path towards faster and sharper Vision Transformers. The source code is available at \href{https://github.com/LeapLabTHU/LinearDiff}{https://github.com/LeapLabTHU/LinearDiff}. Yifan Pu, Jixuan Ying, Qixiu Li, Tianzhu Ye, Dongchen Han, Xinyu Shao, Gao Huang 0001, Xiu Li 0001 |
NeurIPS | 5 |
| 2025 | Feature aggregation and connectivity for object re-identification
Dongchen Han, Baodi Liu, Shuai Shao 0006, Weifeng Liu 0001, Yicong Zhou |
Pattern Recognit. | 1 |
| 2024 | GSVA: Generalized Segmentation via Multimodal Large Language ModelsabstractGeneralized Referring Expression Segmentation (GRES) extends the scope of classic RES to refer to multiple ob-jects in one expression or identify the empty targets absent in the image. GRES poses challenges in modeling the com-plex spatial relationships of the instances in the image and identifying non-existing referents. Multimodal Large Language Models (MLLMs) have recently shown tremendous progress in these complicated vision-language tasks. Con-necting Large Language Models (LLMs) and vision models, MLLMs are proficient in understanding contexts with visual inputs. Among them, LISA, as a representative, adopts a special [SEG] token to prompt a segmentation mask de-coder, e.g., SAM, to enable MLLMs in the RES task. How-ever, existing solutions to GRES remain unsatisfactory since current segmentation MLLMs cannot correctly handle the cases where users might reference multiple subjects in a singular prompt or provide descriptions incongruent with any image target. In this paper, we propose Generalized Segmentation Vision Assistant (GSVA) to address this gap. Specifically, GSVA reuses the [SEG] token to prompt the segmentation model towards supporting multiple mask ref-erences simultaneously and innovatively learns to generate a [REJ] token to reject the null targets explicitly. Ex-periments validate GSVA's efficacy in resolving the GRES issue, marking a notable enhancement and setting a new record on the GRES benchmark gRefCOCO dataset. GSVA also proves effective across various classic referring seg-mentation and comprehension tasks. Code is available at https://github.com/LeapLabTHU/GSVA. Zhuofan Xia, Dongchen Han, Yizeng Han, Xuran Pan, Shiji Song, Gao Huang 0001 |
CVPR | 2 |
| 2024 | Agent Attention: On the Integration of Softmax and Linear Attention
Dongchen Han, Tianzhu Ye, Yizeng Han, Zhuofan Xia, Siyuan Pan, Pengfei Wan 0001, Shiji Song, Gao Huang 0001 |
ECCV (50) | 1 |
| 2024 | Efficient Diffusion Transformer with Step-Wise Dynamic Attention Mediators
Yifan Pu, Zhuofan Xia, Dongchen Han, Qixiu Li, Yuhui Yuan, Ji Li 0006, Yizeng Han, Shiji Song, Gao Huang 0001, Xiu Li 0001 |
ECCV (15) | 4 |
| 2024 | Bridging the Divide: Reconsidering Softmax and Linear AttentionabstractWidely adopted in modern Vision Transformer designs, Softmax attention can effectively capture long-range visual information; however, it incurs excessive computational cost when dealing with high-resolution inputs. In contrast, linear attention naturally enjoys linear complexity and has great potential to scale up to higher-resolution images. Nonetheless, the unsatisfactory performance of linear attention greatly limits its practical application in various scenarios. In this paper, we take a step forward to close the gap between the linear and Softmax attention with novel theoretical analyses, which demystify the core factors behind the performance deviations. Specifically, we present two key perspectives to understand and alleviate the limitations of linear attention: the injective property and the local modeling ability. Firstly, we prove that linear attention is not injective, which is prone to assign identical attention weights to different query vectors, thus adding to severe semantic confusion since different queries correspond to the same outputs. Secondly, we confirm that effective local modeling is essential for the success of Softmax attention, in which linear attention falls short. The aforementioned two fundamental differences significantly contribute to the disparities between these two attention paradigms, which is demonstrated by our substantial empirical validation in the paper. In addition, more experiment results indicate that linear attention, as long as endowed with these two properties, can outperform Softmax attention across various tasks while maintaining lower computation complexity. Code is available at https://github.com/LeapLabTHU/InLine. Dongchen Han, Yifan Pu, Zhuofan Xia, Yizeng Han, Xuran Pan, Xiu Li 0001, Jiwen Lu, Shiji Song, Gao Huang 0001 |
NeurIPS | 1 |
| 2024 | Demystify Mamba in Vision: A Linear Attention PerspectiveabstractMamba is an effective state space model with linear computation complexity. It has recently shown impressive efficiency in dealing with high-resolution inputs across various vision tasks. In this paper, we reveal that the powerful Mamba model shares surprising similarities with linear attention Transformer, which typically underperform conventional Transformer in practice. By exploring the similarities and disparities between the effective Mamba and subpar linear attention Transformer, we provide comprehensive analyses to demystify the key factors behind Mamba’s success. Specifically, we reformulate the selective state space model and linear attention within a unified formulation, rephrasing Mamba as a variant of linear attention Transformer with six major distinctions: input gate, forget gate, shortcut, no attention normalization, single-head, and modified block design. For each design, we meticulously analyze its pros and cons, and empirically evaluate its impact on model performance in vision tasks. Interestingly, the results highlight the forget gate and block design as the core contributors to Mamba’s success, while the other four designs are less crucial. Based on these findings, we propose a Mamba-
Inspired Linear Attention (MILA) model by incorporating the merits of these two key designs into linear attention. The resulting model outperforms various vision Mamba models in both image classification and high-resolution dense prediction tasks, while enjoying parallelizable computation and fast inference speed. Code is available at https://github.com/LeapLabTHU/MLLA. Dongchen Han, Zhuofan Xia, Yizeng Han, Yifan Pu, Chunjiang Ge, Shiji Song, Bo Zheng 0007, Gao Huang 0001 |
NeurIPS | 1 |
| 2023 | Dynamic Perceiver for Efficient Visual RecognitionabstractEarly exiting has become a promising approach to improving the inference efficiency of deep networks. By structuring models with multiple classifiers (exits), predictions for "easy" samples can be generated at earlier exits, negating the need for executing deeper layers. Current multi-exit networks typically implement linear classifiers at intermediate layers, compelling low-level features to encapsulate high-level semantics. This sub-optimal design invariably undermines the performance of later exits. In this paper, we propose Dynamic Perceiver (Dyn-Perceiver) to decouple the feature extraction procedure and the early classification task with a novel dual-branch architecture. A feature branch serves to extract image features, while a classification branch processes a latent code assigned for classification tasks. Bi-directional cross-attention layers are established to progressively fuse the information of both branches. Early exits are placed exclusively within the classification branch, thus eliminating the need for linear separability in low-level features. Dyn-Perceiver constitutes a versatile and adaptable framework that can be built upon various architectures. Experiments on image classification, action recognition, and object detection demonstrate that our method significantly improves the inference efficiency of different backbones, outperforming numerous competitive approaches across a broad range of computational budgets. Evaluation on both CPU and GPU platforms substantiate the superior practical efficiency of Dyn-Perceiver. Code is available at https://www.github.com/LeapLabTHU/Dynamic_Perceiver. Yizeng Han, Dongchen Han, Yulin Wang 0002, Xuran Pan, Yifan Pu, Chao Deng 0002, Junlan Feng, Shiji Song, Gao Huang 0001 |
ICCV | 2 |
| 2023 | FLatten Transformer: Vision Transformer using Focused Linear AttentionabstractThe quadratic computation complexity of self-attention has been a persistent challenge when applying Transformer models to vision tasks. Linear attention, on the other hand, offers a much more efficient alternative with its linear complexity by approximating the Softmax operation through carefully designed mapping functions. However, current linear attention approaches either suffer from significant performance degradation or introduce additional computation overhead from the mapping functions. In this paper, we propose a novel Focused Linear Attention module to achieve both high efficiency and expressiveness. Specifically, we first analyze the factors contributing to the performance degradation of linear attention from two perspectives: the focus ability and feature diversity. To overcome these limitations, we introduce a simple yet effective mapping function and an efficient rank restoration module to enhance the expressiveness of self-attention while maintaining low computation complexity. Extensive experiments show that our linear attention module is applicable to a variety of advanced vision Transformers, and achieves consistently improved performances on multiple benchmarks. Code is available at https://github.com/LeapLabTHU/FLatten-Transformer. Dongchen Han, Xuran Pan, Yizeng Han, Shiji Song, Gao Huang 0001 |
ICCV | 1 |
| 2023 | Non-Contrastive Nearest Neighbor Identity-Guided Method for Unsupervised Object Re-IdentificationabstractRecently, self-paced contrastive learning has emerged as a promising method for unsupervised object re-identification. These methods generate pseudo labels, store centroid features in the memory bank, and periodically update them. However, affected by the performance of the clustering method, within each cluster exists inevitably noisy instances, and self-paced contrastive learning usually requires a large number of negative samples from various classes, where false-negative samples give rise to the class collision issue. These lead to performing incorrect model optimization. In this paper, we propose a non-contrastive nearest neighbor identity-guided (NNNI) method to overcome these challenges. The advantage of NNNI is to provide the model with a highly accurate prior. Specifically, this method relies on the random identity sampler commonly used in re-identification tasks to provide the network with a regression target of the nearest neighbors of the same identity within a mini-batch. It encodes more and more information in an iterative process through a Siamese network with an exponential moving average to train high-quality representations. NNNI alleviates the negative effects of noise instances and corrects class collision issues during training. Extensive experiments show that our method is effective on unsupervised object re-identification and achieves state-of-the-art performance on three large-scale person re-identification datasets and one large-scale vehicle re-identification dataset, which is competitive with even supervised methods. Dongchen Han, Weifeng Liu 0001, Mingchen Zou, Baodi Liu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Pseudo-Q: Generating Pseudo Language Queries for Visual GroundingabstractVisual grounding, i.e., localizing objects in images ac-cording to natural language queries, is an important topic in visual language understanding. The most effective approaches for this task are based on deep learning, which generally require expensive manually labeled image-query or patch-query pairs. To eliminate the heavy depen-dence on human annotations, we present a novel method, named Pseudo-Q, to automatically generate pseudo language queries for supervised training. Our method lever-ages an off-the-shelf object detector to identify visual ob-jects from unlabeled images, and then language queries for these objects are obtained in an unsupervised fashion with a pseudo-query generation module. Then, we design a task-related query prompt module to specifically tailor generated pseudo language queries for visual grounding tasks. Further, in order to fully capture the contextual re-lationships between images and language queries, we de-velop a visual-language model equipped with multi-level cross-modality attention mechanism. Extensive experimen-tal results demonstrate that our method has two notable benefits: (1) it can reduce human annotation costs signifi-cantly, e.g., 31% on Ref Coco [65] without degrading orig-inal model's performance under the fully supervised set-ting, and (2) without bells and whistles, it achieves supe-rior or comparable performance compared to state-of-the-art weakly-supervised visual grounding methods on all the five datasets we have experimented. Code is available at https://github.com/LeapLabTHU/Pseudo-Q. Haojun Jiang, Yuanze Lin, Dongchen Han, Shiji Song, Gao Huang 0001 |
CVPR | 3 |
| 2022 | Contrastive Language-Image Pre-Training with Knowledge GraphsabstractRecent years have witnessed the fast development of large-scale pre-training frameworks that can extract multi-modal representations in a unified form and achieve promising performances when transferred to downstream tasks. Nevertheless, existing approaches mainly focus on pre-training with simple image-text pairs, while neglecting the semantic connections between concepts from different modalities. In this paper, we propose a knowledge-based pre-training framework, dubbed Knowledge-CLIP, which injects semantic information into the widely used CLIP model. Through introducing knowledge-based objectives in the pre-training process and utilizing different types of knowledge graphs as training data, our model can semantically align the representations in vision and language with higher quality, and enhance the reasoning ability across scenarios and modalities. Extensive experiments on various vision-language downstream tasks demonstrate the effectiveness of Knowledge-CLIP compared with the original CLIP and competitive baselines. Xuran Pan, Tianzhu Ye, Dongchen Han, Shiji Song, Gao Huang 0001 |
NeurIPS | 3 |
| 2022 | Object re-identification with distribution corrected ranking list
Dongchen Han, Shuai Shao 0006, Weifeng Liu 0001, Baodi Liu |
Neurocomputing | 1 |
| 2020 | Designed Interactive Toys for Children with Cerebral PalsyabstractChildren with cerebral palsy (CP) need to go through intensive rehabilitation exercises to develop and enhance their fine motor control in daily living. However, most of them cannot persist the regular repetitive exercise session using traditional tools for a long time. To provide a playful and attractive rehabilitation environment, toys are introduced to motivate children for exercises. This study aims to develop diverse toy modules and combine with basic blocks from LEGO to support various hand and arm functional training. The joyful color, cartoon animals, visual and audio feedbacks are proposed to increase the modules' attractiveness. Their interchangeable handles and knobs can support different levels of exercises, which improve the toy modules' accessibility for children with CP. Yixuan Bian, Dongchen Han |
TEI | 3 |