VLDB 2026 Research / reviewers in the wild / expert
Ao Li 0004
dblp:54/2788-4
· DBLP profile ↗
12ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0002-1927-8606ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Efficient and distributed learning · 43% Representation and self-supervised learning · 28% Generative modeling · 15% |
Topics — the 6 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | MaRS: A Fast Sampler for Mean Reverting Diffusion based on ODE and SDE Solvers · ICLR 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.9 | 1 | 2025 | Mutual Effort for Efficiency: A Similarity-based Token Pruning for Vision Transformers in Self-Supervised Learning · ICLR 2025 |
Machine learning › Reinforcement learning
sample efficiency |
0.9 | 1 | 2025 | MaRS: A Fast Sampler for Mean Reverting Diffusion based on ODE and SDE Solvers · ICLR 2025 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › self-supervised visual representation learning
self-supervised vision transformer |
0.9 | 1 | 2025 | Mutual Effort for Efficiency: A Similarity-based Token Pruning for Vision Transformers in Self-Supervised Learning · ICLR 2025 |
Machine learning › Efficient and distributed learning › model compression
token pruning |
0.9 | 1 | 2025 | Mutual Effort for Efficiency: A Similarity-based Token Pruning for Vision Transformers in Self-Supervised Learning · ICLR 2025 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
efficient self-supervised learning |
0.8 | 1 | 2024 | Waxing-and-Waning: a Generic Similarity-based Framework for Efficient Self-Supervised Learning · ICLR 2024 |
Methods — techniques the papers use, named apart from their topics
token pruning · 0.9difficulty-aware pruning · 0.9cross-branch similarity · 0.9SDE solver · 0.9ODE solver · 0.9similarity-based pruning · 0.8data augmentation · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rethinking the Potential of Layer Freezing for DNN Training EfficiencyabstractWith the growing scale of deep neural networks and datasets, training has become increasingly expensive. Layer freezing reduces this cost by stopping updates to selected layers, but frozen layers still require forward propagation to generate activations for later layers. Caching these activations as a surrogate dataset can eliminate this redundant computation, but it faces two key challenges: effectively augmenting cached features and reducing the storage overhead of high-dimensional activations. This paper provides the first systematic study of these challenges and proposes practical solutions. We introduce Similarity-Aware Channel Augmentation to preserve accuracy by caching transformation-sensitive channels with limited overhead. We further incorporate lossy compression and design a progressive compression strategy that exploits the higher compressibility of deeper-layer activations. Our method reduces computation cost, memory usage, and training time while maintaining accuracy. Experiments on NVIDIA Orin Edge GPU further demonstrate training acceleration and significant power savings, highlighting its practicality for resource-constrained training. Chence Yang, Ningxi Cheng, Ci Zhang, Qitao Tan, Sheng Li 0019, Ao Li 0004, Xulong Tang, Shaoyi Huang, Jinzhen Wang, Jundong Li, Xiaoming Zhai, Jin Lu 0001, Geng Yuan |
ACM Great Lakes Symposium on VLSI | 7 |
| 2026 | MTS-LOF: Medical Time-Series Representation Learning via Occlusion-Invariant FeaturesabstractMedical time series data are indispensable in healthcare, providing critical insights for disease diagnosis, treatment planning, and patient management. The exponential growth in data complexity, driven by advanced sensor technologies, has presented challenges related to data labeling. Self-supervised learning (SSL) has emerged as a transformative approach to address these challenges, eliminating the need for extensive human annotation. In this study, we introduce a novel framework for Medical Time Series Representation Learning, known as MTS-LOF. MTS-LOF leverages the strengths of Joint-Embedding SSL and Masked Autoencoder (MAE) methods, offering a unique approach to representation learning for medical time series data. By combining these techniques, MTS-LOF enhances the potential of healthcare applications by providing more sophisticated, context-rich representations. Additionally, MTS-LOF employs a multi-masking strategy to facilitate occlusion-invariant feature learning. This approach allows the model to create multiple views of the data by masking portions of it. By minimizing the discrepancy between the representations of these masked patches and the fully visible patches, MTS-LOF learns to capture rich contextual information within medical time series datasets. The results of experiments conducted on diverse medical time series datasets demonstrate the superiority of MTS-LOF over other methods. These findings hold promise for significantly enhancing healthcare applications by improving representation learning. Furthermore, our work delves into the integration of Joint-Embedding SSL and MAE techniques, shedding light on the intricate interplay between temporal and structural dependencies in healthcare data. This understanding is crucial, as it allows us to grasp the complexities of healthcare data analysis. Ana S. Carreon-Rascon, Xiwen Chen, Geng Yuan, Ao Li 0004 |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | A Computation and Energy Efficient Hardware Architecture for SSL AccelerationabstractIn Computer Vision (CV), the deployment of Convolutional Neural Networks (CNNs) is often hindered by their substantial computational requirements and large labeled datasets. Self-supervised learning (SSL) serves as an effective approach to reducing the reliance on labeled data with the option of augmentation methods to infer and train CNNs. Excluding irrelevant features accelerates learning and improves optimization. We propose a Field-Programmable Gate Array (FPGA)-based hardware accelerator architecture tailored for SSL framework, leveraging its parallelism and reconfigurability to expedite block matching, optimize sparse convolutions, and manage data reuse, significantly improving resource and energy efficiency. The implementation and evaluation of our work on Xilinx ZCU102 FPGA working at 200 MHz confirm that the similarity finding part's FPGA accelerations with a low hardware overhead generates a latency of 0.0106 seconds, surpassing GPU and CPU, and in the sparse CNN's FPGA acceleration part, with the processing of VGG16 and ResNet50, compared with the related FPGA-based works, our design claims a maximum of 3.08× throughput improvement and 1.5× in energy efficiency. Huidong Ji, Sheng Li 0019, Chen Ding 0010, Jiawei Xu 0001, Qitao Tan, Jun Liu 0075, Ao Li 0004, Xulong Tang, Lirong Zheng 0001, Geng Yuan, Zhuo Zou |
ASP-DAC | 8 |
| 2025 | AngioDiff: Structure-Preserving and 3D-Consistent Diffusion for CT Angiography SynthesisabstractComputed tomography angiography (CTA) plays a crucial role in the diagnosis of thoracic vascular diseases, yet the need for iodinated contrast agents raises safety and accessibility concerns. Generating synthetic contrast-enhanced CT (CECT) from non-contrast CT (NCCT) offers a promising solution to these challenges. However, existing generative methods often struggle to preserve fine vascular details and continuity along the anisotropic z-axis in volumetric data. In this work, we present AngioDiff, a novel diffusion-based framework for high-fidelity CTA synthesis from NCCT. Our method leverages a conditional diffusion model with a mean reverting prior and incorporates a sliding window mechanism with asynchronous denoising to effectively model large-scale 3D CT volumes. To further stream-line the process and address edge artifacts, we introduce a sequence padding strategy that simplifies training and sampling while enhancing structural continuity. Furthermore, our network design combines spatial and axial attention modules to adequately capture intra-slice and inter-slice dependencies. Comprehensive experiments on multi-center datasets demonstrate that AngioDiff consistently outperforms state-of-the-art methods in both 2D slice-based and 3D volume-based quantitative metrics, achieving superior anatomical fidelity and volumetric consistency. This work highlights the clinical potential of diffusion-based models for agent-free CTA, offering a safer and more accessible alter-native to traditional imaging protocols. Ao Li 0004, Wei Fang 0005, Ge Yang 0002, Minfeng Xu |
BIBM | 1 |
| 2025 | Towards Memory-Efficient and Sustainable Machine Unlearning on Edge using Zeroth-Order Optimizer
Ci Zhang, Chence Yang, Qitao Tan, Jun Liu 0075, Ao Li 0004, Yanzhi Wang 0001, Jin Lu 0001, Geng Yuan |
ACM Great Lakes Symposium on VLSI | 5 |
| 2025 | Mutual Effort for Efficiency: A Similarity-based Token Pruning for Vision Transformers in Self-Supervised LearningabstractSelf-supervised learning (SSL) offers a compelling solution to the challenge of extensive labeled data requirements in traditional supervised learning.
With the proven success of Vision Transformers (ViTs) in supervised tasks, there is increasing interest in adapting them for SSL frameworks. However, the high computational demands of SSL pose substantial challenges, particularly on resource-limited platforms like edge devices, despite its ability to achieve high accuracy without labeled data.
Recent studies in supervised learning have shown that token pruning can reduce training costs by removing less informative tokens without compromising accuracy. However, SSL’s dual-branch encoders make traditional single-branch pruning strategies less effective, as they fail to account for the critical cross-branch similarity information, leading to reduced accuracy in SSL.
To this end, we introduce SimPrune, a novel token pruning strategy designed for ViTs in SSL. SimPrune leverages cross-branch similarity information to efficiently prune tokens, retaining essential semantic information across dual branches. Additionally, we incorporate a difficulty-aware pruning strategy to further enhance SimPrune's effectiveness.
Experimental results show that our proposed approach effectively reduces training computation while maintaining accuracy. Specifically, our approach offers 24\% savings in training costs compared to SSL baseline, without sacrificing accuracy. Sheng Li 0019, Qitao Tan, Yue Dai 0005, Zhenglun Kong, Jun Liu 0075, Ao Li 0004, Ninghao Liu 0001, Yufei Ding 0001, Xulong Tang, Geng Yuan |
ICLR | 7 |
| 2025 | MaRS: A Fast Sampler for Mean Reverting Diffusion based on ODE and SDE SolversabstractIn applications of diffusion models, controllable generation is of practical significance, but is also challenging. Current methods for controllable generation primarily focus on modifying the score function of diffusion models, while Mean Reverting (MR) Diffusion directly modifies the structure of the stochastic differential equation (SDE), making the incorporation of image conditions simpler and more natural. However, current training-free fast samplers are not directly applicable to MR Diffusion. And thus MR Diffusion requires hundreds of NFEs (number of function evaluations) to obtain high-quality samples. In this paper, we propose a new algorithm named MaRS (MR Sampler) to reduce the sampling NFEs of MR Diffusion. We solve the reverse-time SDE and the probability flow ordinary differential equation (PF-ODE) associated with MR Diffusion, and derive semi-analytical solutions. The solutions consist of an analytical function and an integral parameterized by a neural network. Based on this solution, we can generate high-quality samples in fewer steps. Our approach does not require training and supports all mainstream parameterizations, including noise prediction, data prediction and velocity prediction. Extensive experiments demonstrate that MR Sampler maintains high sampling quality with a speedup of 10 to 20 times across ten different image restoration tasks. Our algorithm accelerates the sampling procedure of MR Diffusion, making it more practical in controllable generation. Ao Li 0004, Wei Fang 0005, Le Lu 0001, Ge Yang 0002, Minfeng Xu |
ICLR | 1 |
| 2024 | Waxing-and-Waning: a Generic Similarity-based Framework for Efficient Self-Supervised LearningabstractDeep Neural Networks (DNNs), essential for diverse applications such as visual recognition and eldercare, often require a large amount of labeled data for training, making widespread deployment of DNNs a challenging task. Self-supervised learning (SSL) emerges as a promising approach, which leverages inherent patterns within data through diverse augmentations to train models without explicit labels. However, while SSL has shown notable advancements in accuracy, its high computation costs remain a daunting impediment, particularly for resource-constrained platforms. To address this problem, we introduce SimWnW, a similarity-based efficient self-supervised learning framework. By strategically removing less important regions in augmented images and feature maps, SimWnW not only reduces computation costs but also eliminates irrelevant features that might slow down the learning process, thereby accelerating model convergence. The experimental results show that SimWnW effectively reduces the amount of computation costs in self-supervised model training without compromising accuracy. Specifically, SimWnW yields up to 54\% and 51\% computation savings in training from scratch and transfer learning tasks, respectively. Sheng Li 0019, Chao Wu 0006, Ao Li 0004, Yanzhi Wang 0001, Xulong Tang, Geng Yuan |
ICLR | 3 |
| 2024 | Knowledge distillation under ideal joint classifier assumption
Xiwen Chen, Gregory Ditzler, Janet Roveda, Ao Li 0004 |
Neural Networks | 5 |
| 2024 | DeScoD-ECG: Deep Score-Based Diffusion Model for ECG Baseline Wander and Noise RemovalabstractOBJECTIVE: Electrocardiogram (ECG) signals commonly suffer noise interference, such as baseline wander. High-quality and high-fidelity reconstruction of the ECG signals is of great significance to diagnosing cardiovascular diseases. Therefore, this paper proposes a novel ECG baseline wander and noise removal technology. METHODS: We extended the diffusion model in a conditional manner that was specific to the ECG signals, namely the Deep Score-Based Diffusion model for Electrocardiogram baseline wander and noise removal (DeScoD-ECG). Moreover, we deployed a multi-shots averaging strategy that improved signal reconstructions. We conducted the experiments on the QT Database and the MIT-BIH Noise Stress Test Database to verify the feasibility of the proposed method. Baseline methods are adopted for comparison, including traditional digital filter-based and deep learning-based methods. RESULTS: The quantities evaluation results show that the proposed method obtained outstanding performance on four distance-based similarity metrics with at least 20% overall improvement compared with the best baseline method. CONCLUSION: This paper demonstrates the state-of-the-art performance of the DeScoD-ECG for ECG baseline wander and noise removal, which has better approximations of the true data distribution and higher stability under extreme noise corruptions. SIGNIFICANCE: This study is one of the first to extend the conditional diffusion-based generative model for ECG noise removal, and the DeScoD-ECG has the potential to be widely used in biomedical applications. Gregory Ditzler, Janet Roveda, Ao Li 0004 |
IEEE J. Biomed. Health Informatics | 4 |
| 2020 | Sequence-level Supervised Deep Neural Networks for Mitosis Event Detection in Time-Lapse Microscopy ImagesabstractAutomatic mitosis detection is a key step in measuring cell proliferation and analyzing the responses to various stimuli. Current deep neural networks can learn complex visual features and capture long-range temporal dependencies. However, the state-of-the-art mitosis detection models require massive ground truth annotations which is labor intensive in biomedical experiments. Therefore, we propose a sequence-level supervised neural networks model to detect mitosis events at pixel-and-frame level. By using binary labels, the proposed network is trained to predict the presence of mitosis for the input microscopy sequences. Then we leverage the feature map produced by the proposed network to localize the cell division. The proposed model achieved a detection F1-score 0.881.With significantly less amount of ground truth in the training data, our method achieved competitive performance compared with the state-of-art fully supervised mitosis detection methods. Siteng Chen, Ao Li 0004, Janet Roveda |
BIBM | 2 |
| 2019 | Weakly Supervised Deep Learning for Detecting and Counting Dead Cells in Microscopy ImagesabstractCounting dead cells is a key step in evaluating the performance of chemotherapy treatment and drug screening. Deep convolutional neural networks (CNNs) can learn complex visual features, but require massive ground truth annotations which is expensive in biomedical experiments. Counting cells, especially dead cells, with very few ground truth annotations remains unexplored. In this paper, we automate dead cell counting using a weakly supervised strategy. We took advantage of the fact that cell death is low before chemotherapy treatment and increases after treatment. Motivated by the contrast, we first design image level supervised only classification neural networks to detect dead cells. Based on the class response map in classification networks, we calculate a Dead Confidence Map (DCM) to specify confidence of each dead cell. Associated with peak clustering, local maximums in the DCM are used to count the number of dead cells. In addition, a biological experiment based weakly supervised data preparation strategy is proposed to minimize human intervention. We show classification performance compared to general purpose and cell classification networks, and report results for the image-level supervised counting task. Siteng Chen, Ao Li 0004, Kathleen Lasick, Julie Huynh, Linda S. Powers, Janet Roveda, Andrew Paek |
ICMLA | 2 |