Ao Li 0007

dblp:54/2788-7 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0002-8866-2302ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 4 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Efficient and distributed learning · 78% Segmentation and scene understanding · 16% Trustworthy machine learning · 5%
Computer graphics and multimedia
4 papers
Image and video processing · 100%
Network and information security
1 paper
Hardware security and side channels · 100%

Topics — the 25 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
1.722025
Exploring Frequency-Inspired Optimization in Transformer for Efficient Single Image Super-Resolution · IEEE Trans. Pattern Anal. Mach. Intell. 2025
MobileIE: An Extremely Lightweight and Effective ConvNet for Real-Time Image Enhancement on Mobile Devices · ICCV 2025
Image and video processing › super-resolution
image super-resolution
1.522025
Exploring Frequency-Inspired Optimization in Transformer for Efficient Single Image Super-Resolution · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Feature Modulation Transformer: Cross-Refinement of Global Representation via High-Frequency Prior for Image Super-Resolution · ICCV 2023
Image and video processing › super-resolution › image super-resolution
single image super-resolution
1.522025
Exploring Frequency-Inspired Optimization in Transformer for Efficient Single Image Super-Resolution · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Feature Modulation Transformer: Cross-Refinement of Global Representation via High-Frequency Prior for Image Super-Resolution · ICCV 2023
Computer vision › Segmentation and scene understanding › instance segmentation
camouflaged object segmentation
0.912025
Towards Real Zero-Shot Camouflaged Object Segmentation Without Camouflaged Annotations · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Machine learning › Efficient and distributed learning › federated learning
contribution evaluation
0.912025
Subspace Constraint and Contribution Estimation for Heterogeneous Federated Learning · CVPR 2025
Machine learning › Efficient and distributed learning
federated learning
0.912025
Subspace Constraint and Contribution Estimation for Heterogeneous Federated Learning · CVPR 2025
Machine learning › Efficient and distributed learning › federated learning
heterogeneous federated learning
0.912025
Subspace Constraint and Contribution Estimation for Heterogeneous Federated Learning · CVPR 2025
Machine learning › Efficient and distributed learning › model compression
lightweight neural network
0.912025
MobileIE: An Extremely Lightweight and Effective ConvNet for Real-Time Image Enhancement on Mobile Devices · ICCV 2025
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.912025
Rethinking Token Reduction with Parameter-Efficient Fine-Tuning in ViT for Pixel-Level Tasks · CVPR 2025
Machine learning › Efficient and distributed learning › model compression › quantization
post-training quantization
0.912025
Exploring Frequency-Inspired Optimization in Transformer for Efficient Single Image Super-Resolution · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Machine learning › Efficient and distributed learning
token reduction
0.912025
Rethinking Token Reduction with Parameter-Efficient Fine-Tuning in ViT for Pixel-Level Tasks · CVPR 2025
Image and video processing
image enhancement
0.912025
MobileIE: An Extremely Lightweight and Effective ConvNet for Real-Time Image Enhancement on Mobile Devices · ICCV 2025
Image and video processing
image restoration
0.912025
Deep-Learning-Empowered Super Resolution: A Comprehensive Survey and Future Prospects · Proc. IEEE 2025
Image and video processing › super-resolution
learning-based super-resolution
0.912025
Deep-Learning-Empowered Super Resolution: A Comprehensive Survey and Future Prospects · Proc. IEEE 2025
Image and video processing › image enhancement
real-time image enhancement
0.912025
MobileIE: An Extremely Lightweight and Effective ConvNet for Real-Time Image Enhancement on Mobile Devices · ICCV 2025
Image and video processing
super-resolution
0.912025
Deep-Learning-Empowered Super Resolution: A Comprehensive Survey and Future Prospects · Proc. IEEE 2025
Hardware security and side channels › side-channel attack
profiled side-channel attack
0.912025
Enhancing Model Generalization for Efficient Cross-Device Side-Channel Analysis · IEEE Trans. Inf. Forensics Secur. 2025
Hardware security and side channels
side-channel attack
0.912025
Enhancing Model Generalization for Efficient Cross-Device Side-Channel Analysis · IEEE Trans. Inf. Forensics Secur. 2025
Machine learning › Trustworthy machine learning › robustness
overfitting mitigation
0.312025
Subspace Constraint and Contribution Estimation for Heterogeneous Federated Learning · CVPR 2025
Computer vision › Segmentation and scene understanding › medical image segmentation
polyp segmentation
0.312025
Towards Real Zero-Shot Camouflaged Object Segmentation Without Camouflaged Annotations · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Machine learning › Trustworthy machine learning
robustness
0.312025
Subspace Constraint and Contribution Estimation for Heterogeneous Federated Learning · CVPR 2025
Computer vision › Segmentation and scene understanding › saliency detection
salient object detection
0.312025
Towards Real Zero-Shot Camouflaged Object Segmentation Without Camouflaged Annotations · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Image and video processing › super-resolution
video super-resolution
0.312025
Deep-Learning-Empowered Super Resolution: A Comprehensive Survey and Future Prospects · Proc. IEEE 2025
Hardware security and side channels › side-channel attack › profiled side-channel attack
deep learning-based side-channel attack
0.312025
Enhancing Model Generalization for Efficient Cross-Device Side-Channel Analysis · IEEE Trans. Inf. Forensics Secur. 2025
Embedded and real-time systems
mobile computing
0.312025
MobileIE: An Extremely Lightweight and Effective ConvNet for Real-Time Image Enhancement on Mobile Devices · ICCV 2025

Methods — techniques the papers use, named apart from their topics

knowledge distillation · 3.5reparameterization · 2.6attention mechanism · 2.6transformer · 1.5vision transformer · 0.9token reduction · 0.9subspace constraint · 0.9parameter-space distance · 0.9multi-task loss · 0.9masked image modeling · 0.9knowledge transfer via auxiliary models · 0.9generative adversarial network · 0.9frequency-guided quantization · 0.9feature-space distance · 0.9denoising diffusion probabilistic model · 0.9deep learning · 0.9convolutional neural network · 0.9adaptive dual clipping · 0.9
YearPublicationVenuePosition
2025 Rethinking Token Reduction with Parameter-Efficient Fine-Tuning in ViT for Pixel-Level Tasks
abstract
Parameter-Efficient fine-tuning (PEFT) adapts pre-trained models to new tasks by updating only a small subset of parameters, achieving efficiency but still facing significant inference costs driven by input token length. This challenge is even more pronounced in pixel-level tasks, which require longer input sequences compared to image-level tasks. Although token reduction (TR) techniques can help reduce computational demands, they often lead to homogeneous attention patterns that compromise performance in pixel-level scenarios. This study underscores the importance of maintaining attention diversity for these tasks and proposes to enhance attention diversity while ensuring the completeness of token sequences. Our approach effectively reduces the number of tokens processed within transformer blocks, improving computational efficiency without sacrificing performance on several pixel-level tasks. We also demonstrate the superior generalization capability of our proposed method compared to challenging baseline models. The source code will be made available at https://github.com/AVC2-UESTC/DAR-TR-PEFT.
Ao Li 0007, Hu Yao, Ce Zhu, Le Zhang 0001
CVPR2
2025 Subspace Constraint and Contribution Estimation for Heterogeneous Federated Learning
abstract
Heterogeneous Federated Learning (HFL) has received widespread attention due to its adaptability to different models and data. The HFL approach utilizing auxiliary models for knowledge transfer can further enhance flexibility. However, existing frameworks face the challenges of local overfitting and aggregation bias. To address these issues, we propose FedSCE. By restricting specific layers of the local model updates to a subspace, FedSCE reduces the degrees of freedom of the update, enhances generalization, and mitigates the risk of overfitting. The subspace is dynamically updated to ensure coverage of the latest model update trajectory. Additionally, FedSCE evaluates client contributions based on the update distance of the auxiliary model in feature space and parameter space, achieving adaptive weighted aggregation. We validate our approach in both feature-skewed and label-skewed scenarios, demonstrating that on Office10, our method exceeds the best baseline by 3.87%. The code will be available at https://github.com/AVC2-UESTC/FedSCE.git.
Xiangtao Zhang, Ao Li 0007, Yipeng Liu 0001, Fan Zhang 0013, Ce Zhu, Le Zhang 0001
CVPR3
2025 MobileIE: An Extremely Lightweight and Effective ConvNet for Real-Time Image Enhancement on Mobile Devices
abstract
Recent advancements in deep neural networks have driven significant progress in image enhancement (IE). However, deploying deep learning models on resource-constrained platforms, such as mobile devices, remains challenging due to high computation and memory demands. To address these challenges and facilitate real-time IE on mobile, we introduce an extremely lightweight Convolutional Neural Network (CNN) framework with around 4K parameters. Our approach integrates reparameterization with an Incremental Weight Optimization strategy to ensure efficiency. Additionally, we enhance performance with a Feature Self-Transform module and a Hierarchical Dual-Path Attention mechanism, optimized with a Local Variance-Weighted loss. With this efficient framework, we are the first to achieve real-time IE inference at up to 1,100 frames per second (FPS) while delivering competitive image quality, achieving the best trade-off between speed and performance across multiple IE tasks. The code will be available at https://github.com/AVC2-UESTC/MobileIE.git.
Hailong Yan, Ao Li 0007, Xiangtao Zhang, Zhe Liu 0019, Zenglin Shi, Ce Zhu, Le Zhang 0001
ICCV2
2025 Towards Real Zero-Shot Camouflaged Object Segmentation Without Camouflaged Annotations
abstract
Camouflaged Object Segmentation (COS) faces significant challenges due to the scarcity of annotated data, where meticulous pixel-level annotation is both labor-intensive and costly, primarily due to the intricate object-background boundaries. Addressing the core question, "Can COS be effectively achieved in a zero-shot manner without manual annotations for any camouflaged object?", we propose an affirmative solution. We examine the learned attention patterns for camouflaged objects and introduce a robust zero-shot COS framework. Our findings reveal that while transformer models for salient object segmentation (SOS) prioritize global features in their attention mechanisms, camouflaged object segmentation exhibits both global and local attention biases. Based on these findings, we design a framework that adapts with the inherent local pattern bias of COS while incorporating global attention patterns and a broad semantic feature space derived from SOS. This enables efficient zero-shot transfer for COS. Specifically, We incorporate a Masked Image Modeling (MIM) based image encoder optimized for Parameter-Efficient Fine-Tuning (PEFT), a Multimodal Large Language Model (M-LLM), and a Multi-scale Fine-grained Alignment (MFA) mechanism. The MIM encoder captures essential local features, while the PEFT module learns global and semantic representations from SOS datasets. To further enhance semantic granularity, we leverage the M-LLM to generate caption embeddings conditioned on visual cues, which are meticulously aligned with multi-scale visual features via MFA. This alignment enables precise interpretation of complex semantic contexts. Moreover, we introduce a learnable codebook to represent the M-LLM during inference, significantly reducing computational demands while maintaining performance. Our framework demonstrates its versatility and efficacy through rigorous experimentation, achieving state-of-the-art performance in zero-shot COS with $F_{\beta }^{w}$Fβw scores of 72.9% on CAMO and 71.7% on COD10K. By removing the M-LLM during inference, we achieve an inference speed comparable to that of traditional end-to-end models, reaching 18.1 FPS. Additionally, our method excels in polyp segmentation, and underwater scene segmentation, outperforming challenging baselines in both zero-shot and supervised settings, thereby implying its potentiality in various segmentation tasks.
Tian-Zhu Xiang, Ao Li 0007, Ce Zhu, Le Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Exploring Frequency-Inspired Optimization in Transformer for Efficient Single Image Super-Resolution
abstract
Transformer-based methods have exhibited remarkable potential in single image super-resolution (SISR) by effectively extracting long-range dependencies. However, most of the current research in this area has prioritized the design of transformer blocks to capture global information, while overlooking the importance of incorporating high-frequency priors, which we believe could be beneficial. In our study, we conducted a series of experiments and found that transformer structures are more adept at capturing low-frequency information, but have limited capacity in constructing high-frequency representations when compared to their convolutional counterparts. Our proposed solution, the cross-refinement adaptive feature modulation transformer (CRAFT), integrates the strengths of both convolutional and transformer structures. It comprises three key components: the high-frequency enhancement residual block (HFERB) for extracting high-frequency information, the shift rectangle window attention block (SRWAB) for capturing global information, and the hybrid fusion block (HFB) for refining the global representation. To tackle the inherent intricacies of transformer structures, we introduce a frequency-guided post-training quantization (PTQ) method aimed at enhancing CRAFT's efficiency. These strategies incorporate adaptive dual clipping and boundary refinement. To further amplify the versatility of our proposed approach, we extend our PTQ strategy to function as a general quantization method for transformer-based SISR techniques. Our experimental findings showcase CRAFT's superiority over current state-of-the-art methods, both in full-precision and quantization scenarios. These results underscore the efficacy and universality of our PTQ strategy.
Ao Li 0007, Le Zhang 0001, Yun Liu 0011, Ce Zhu
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Deep-Learning-Empowered Super Resolution: A Comprehensive Survey and Future Prospects
abstract
Super-resolution (SR) has garnered significant attention within the computer vision community, driven by advances in deep learning (DL) techniques and the growing demand for high-quality visual applications. With the expansion of this field, numerous surveys have emerged. Most existing surveys focus on specific domains, lacking a comprehensive overview of this field. Here, we present an in-depth review of diverse SR methods, encompassing single-image SR (SISR), video SR (VSR), stereo SR (SSR), and light field SR (LFSR). We extensively cover over 150 SISR methods, nearly 70 VSR approaches, and approximately 30 techniques for SSR and LFSR. We analyze methodologies, datasets, evaluation protocols, empirical results, and complexity. In addition, we conducted a taxonomy based on each backbone structure according to the diverse purposes. We also explore valuable yet understudied open issues in the field. We believe that this work will serve as a valuable resource and offer guidance to researchers in this domain. To facilitate access to related work, we created a dedicated repository available at https://github.com/AVC2-UESTC/Holistic-Super-Resolution-Review
Le Zhang 0001, Ao Li 0007, Qibin Hou, Ce Zhu, Yonina C. Eldar
Proc. IEEE2
2025 Striving for understanding: Deconstructing neural networks in side-channel analysis
Bo Wang 0011, Changshan Su, Ao Li 0007, Gen Li 0011, Yuxing Tang
Pattern Recognit.4
2025 Enhancing Model Generalization for Efficient Cross-Device Side-Channel Analysis
abstract
Deep learning (DL)-based techniques have garnered significant attention as an innovative method for profiled side-channel analysis (SCA). Despite their proven effectiveness, recent studies have highlighted challenges faced by DL-based profiled attacks in a more realistic portability threat model, where two devices are used respectively for profiling and the attack. In this paper, we propose a novel approach for cross-device attack by incorporating the Denoising Diffusion Probabilistic Model (DDPM) to develop a generalized model. Additionally, an adaptive multi-task loss is employed to balance multiple training objectives that respectively focus on model generalization and precision. We evaluate our strategy on five cross-device SCA datasets. The experimental results show that, compared to baseline methods, our approach achieves significantly enhanced performance, as measured by the number of traces required to recover the secret key. Specifically, on a more challenging dataset obtained from three SAKURA-G evaluation boards, our method successfully recovers the secret key using approximately 300 traces, whereas baseline methods fail to guarantee a successful cross-device attack even with 5,000 traces. Furthermore, our method demonstrates remarkably enhanced attack efficiency, reducing attack time by over an hour compared to the baselines.
Bo Wang 0011, Changshan Su, Ao Li 0007, Yuxing Tang, Gen Li 0011
IEEE Trans. Inf. Forensics Secur.4
2023 Feature Modulation Transformer: Cross-Refinement of Global Representation via High-Frequency Prior for Image Super-Resolution
abstract
Transformer-based methods have exhibited remarkable potential in single image super-resolution (SISR) by effectively extracting long-range dependencies. However, most of the current research in this area has prioritized the design of transformer blocks to capture global information, while overlooking the importance of incorporating high-frequency priors, which we believe could be beneficial. In our study, we conducted a series of experiments and found that transformer structures are more adept at capturing low-frequency information, but have limited capacity in constructing high-frequency representations when compared to their convolutional counterparts. Our proposed solution, the cross-refinement adaptive feature modulation transformer (CRAFT), integrates the strengths of both convolutional and transformer structures. It comprises three key components: the high-frequency enhancement residual block (HFERB) for extracting high-frequency information, the shift rectangle window attention block (SRWAB) for capturing global information, and the hybrid fusion block (HFB) for refining the global representation. Our experiments on multiple datasets demonstrate that CRAFT outperforms state-of-the-art methods by up to 0.29dB while using fewer parameters. The source code will be made available at: https://github.com/AVC2-UESTC/CRAFT-SR.git.
Ao Li 0007, Le Zhang 0001, Yun Liu 0011, Ce Zhu
ICCV1
2020 Real-Time Tracking of Vehicles with Siamese Network and Backward Prediction
abstract
Tracking of vehicles is a key technique for Intelligent transportation system, which commonly follows tracking-by-detection strategy. Due to high appearance similarity among vehicles and heavy occlusion caused by busy traffic flow, a major challenge in such a tracking system is the limited performance of the underlying detector which may produce noisy detections. Consequently, Siamese network and backward prediction-based vehicle tracking approach is proposed. Siamese network based forward position prediction is designed to alleviate the interference of noisy detections, while backward prediction verification is performed to reduce the false positives arising with forward prediction. The final tracklets are obtained through weighted merging based on the detection confidence and forward prediction confidence. The experiment results demonstrate that the proposed method outperforms the state-of-the-art on the UA-DETRAC vehicle tracking dataset, as well as maintains real-time processing at an average tracking speed of 20.1fps, which can be used for real-time applications.
Ao Li 0007, Lei Luo 0003, Shu Tang
ICME1
2020 An Improved Puncturing Scheme for Polar Codes
abstract
IoT is widely used and plays an critical role in transforming traditional industries, leading emerging industries, improving people's lives, and providing national security. From the physical layer of communication, it is particularly important to ensure the reliability of transmission in the IoT. Polar code has outstanding performance, but its coding structure determines that the length of polar codes must be the power of 2. The structure of polar code is more flexible with the existing puncturing schemes, but the decoding performance of these schemes varies with the number of puncturing bits. The decoding performance of the puncturing scheme in C0 mode drops sharply when the number of puncturing bits is large, and the decoding performance of the puncturing scheme in C1 mode is poor when the number of puncturing bits is small. In this paper, an improved polar code puncturing scheme is proposed based on the forward sequential puncturing scheme and the bit reversal puncturing scheme in the C0 puncturing mode. Compared with the traditional puncturing scheme, the simulation results indicated that the proposed scheme solves the problem of decoding performance degradation when the number of puncturing bits is too large in C0 mode; compared with the scheme in C1 mode, the scheme in this paper has a decoding performance gain of about 0.2dB when the block error rate reaches 10-3, which can also be achieved in high or low code rate.
Ao Li 0007, Peng Yu 0001, Fanqin Zhou
IWCMC3