Zekun Ai

dblp:355/8074 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0003-1585-2014ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Efficient and distributed learning · 93% Deep learning architectures and training · 7%
Computer graphics and multimedia
2 papers
Image and video processing · 100%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model acceleration
1.522024
SkipVSR: Adaptive Patch Routing for Video Super-Resolution with Inter-Frame Mask · ACM Multimedia 2024
AdaFormer: Efficient Transformer with Adaptive Token Sparsification for Image Super-resolution · AAAI 2024
Machine learning › Efficient and distributed learning › adaptive computation
adaptive computation skipping
0.812024
SkipVSR: Adaptive Patch Routing for Video Super-Resolution with Inter-Frame Mask · ACM Multimedia 2024
Machine learning › Efficient and distributed learning › model compression › token compression
token sparsification
0.812024
AdaFormer: Efficient Transformer with Adaptive Token Sparsification for Image Super-resolution · AAAI 2024
Image and video processing › super-resolution › video super-resolution
efficient video super-resolution
0.812024
SkipVSR: Adaptive Patch Routing for Video Super-Resolution with Inter-Frame Mask · ACM Multimedia 2024
Image and video processing › super-resolution
image super-resolution
0.812024
AdaFormer: Efficient Transformer with Adaptive Token Sparsification for Image Super-resolution · AAAI 2024
Image and video processing › super-resolution › image super-resolution
lightweight super-resolution
0.812024
AdaFormer: Efficient Transformer with Adaptive Token Sparsification for Image Super-resolution · AAAI 2024
Image and video processing › super-resolution
video super-resolution
0.812024
SkipVSR: Adaptive Patch Routing for Video Super-Resolution with Inter-Frame Mask · ACM Multimedia 2024
Machine learning › Deep learning architectures and training
transformer
0.212024
AdaFormer: Efficient Transformer with Adaptive Token Sparsification for Image Super-resolution · AAAI 2024
Image and video processing
video restoration
0.212024
SkipVSR: Adaptive Patch Routing for Video Super-Resolution with Inter-Frame Mask · ACM Multimedia 2024

Methods — techniques the papers use, named apart from their topics

uncertainty-guided loss · 1.5temporal feature alignment · 1.5recurrent network · 1.5early-exit strategy · 1.5confidence estimation · 1.5adaptive token sparsification · 1.5adaptive patch routing · 1.5
YearPublicationVenuePosition
2025 S3SR: Towards Efficient Image Super-Resolution with Selective State Space Model
abstract
Though Transformer-based image super-resolution (SR) has made remarkable progress, the burdensome computation complexity hinders its applications in memory-limited devices. Existing efficient Transformer-based image SR methods mainly focus on designing efficient local window self-attention mechanisms to improve computational efficiency. However, the limited receptive field of local windows often fails to capture global contextual information effectively. Recently, the Selective State Space Model, e.g., Mamba, has shown powerful potential for long-range dependencies modeling with linear complexity. In this work, we propose a selective state space model for efficient image SR, dubbed S3SR. Specifically, we design the Local-then-Global Fusion Block as the core component, which employs different convolution and a 2D cross scan mechanism to take advantage of local patch texture and global relevance. Extensive experiments have demonstrated the superiority of our S3SR, which even outperforms the efficient Transformer-based SR methods, using less computational cost but with a larger global receptive field.
Xiaotong Luo, Zekun Ai, Yanyun Qu
ICME3
2024 AdaFormer: Efficient Transformer with Adaptive Token Sparsification for Image Super-resolution
abstract
Efficient transformer-based models have made remarkable progress in image super-resolution (SR). Most of these works mainly design elaborate structures to accelerate the inference of the transformer, where all feature tokens are propagated equally. However, they ignore the underlying characteristic of image content, i.e., various image regions have distinct restoration difficulties, especially for large images (2K-8K), failing to achieve adaptive inference. In this work, we propose an adaptive token sparsification transformer (AdaFormer) to speed up the model inference for image SR. Specifically, a texture-relevant sparse attention block with parallel global and local branches is introduced, aiming to integrate informative tokens from the global view instead of only in fixed local windows. Then, an early-exit strategy is designed to progressively halt tokens according to the token importance. To estimate the plausibility of each token, we adopt a lightweight confidence estimator, which is constrained by an uncertainty-guided loss to obtain a binary halting mask about the tokens. Experiments on large images have illustrated that our proposal reduces nearly 90% latency against SwinIR on Test8K, while maintaining a comparable performance.
Xiaotong Luo, Zekun Ai, Qiuyuan Liang, Ding Liu 0001, Yuan Xie 0006, Yanyun Qu, Yun Fu 0001
AAAI2
2024 SkipVSR: Adaptive Patch Routing for Video Super-Resolution with Inter-Frame Mask
abstract
Deep neural networks have revealed enormous potential in video super-resolution (VSR), yet the expensive computational expense limits their deployment on resource-limited devices and actual scenarios, especially for restoring multiple frames simultaneously. Existing VSR models contain considerable redundant filters, which drag down the inference efficiency. To accelerate the inference of VSR models, we propose a scalable method based on adaptive patch routing to achieve practical speedup. Specifically, we design a confidence estimator to predict the aggregation performance of each block for adjacent patch information. It learns to dynamically perform block skipping, i.e., choose which basic blocks of the VSR network to execute during inference so as to reduce total computation to the maximum extent without degrading reconstruction accuracy dramatically. However, we observe that skipping error would be amplified as the hidden states propagate along with recurrent networks. To alleviate the issue, we design temporal feature alignment to guarantee the performance. This proposal essentially proposes an adaptive routing scheme for each patch. Extensive experiments demonstrate that our method can not only accelerate inference but also provide strong quantitative and qualitative results. Built upon the BasicVSR model, our method achieves a speedup of 20% on average, going as high as 50% for some images, while even maintaining competitive performance on REDS4.
Zekun Ai, Xiaotong Luo, Yanyun Qu, Yuan Xie 0006
ACM Multimedia1
2023 Joint Feature Aggregation for Stereo Image Super-resolution
abstract
Stereo image super-resolution (Stereo SR) has been a newly rising and challenging problem with the popular application of dual cameras, which can be used to promote the SR performance by adding auxiliary information from another viewpoint. Most of the existing excellent works have concentrated on leveraging the intrinsic feature correlation of two view images via exploring the non-local attention mechanism. However, they only perform interaction once for feature registration and fusion accompanied by the complex view transition constraint, which cannot fully take advantage of the information in the stereo image pairs. In this paper, we propose a joint feature aggregation network for Stereo SR to calibrate single-view features and integrate cross-view knowledge effectively. Specifically, we introduce a self-calibrated feature extractor to excavate multi-scale and multi-direction features within the single-view image. What’s more, we design an adaptive fusion module with the cross-view attention mechanism, which is utilized to mine and fuse the long-range dependencies between the stereo image pairs so as to get rid of the inflexible cycle constraints. Extensive experimental results demonstrate that our proposal successfully achieves superior performance against the state-of-the-art methods on four datasets.
Zekun Ai, Xiaotong Luo, Yanyun Qu
ICME1