EDBT 2026 Demo / reviewers in the wild / expert
Qitao Tan
dblp:281/5609
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0002-1280-0334ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rethinking the Potential of Layer Freezing for DNN Training EfficiencyabstractWith the growing scale of deep neural networks and datasets, training has become increasingly expensive. Layer freezing reduces this cost by stopping updates to selected layers, but frozen layers still require forward propagation to generate activations for later layers. Caching these activations as a surrogate dataset can eliminate this redundant computation, but it faces two key challenges: effectively augmenting cached features and reducing the storage overhead of high-dimensional activations. This paper provides the first systematic study of these challenges and proposes practical solutions. We introduce Similarity-Aware Channel Augmentation to preserve accuracy by caching transformation-sensitive channels with limited overhead. We further incorporate lossy compression and design a progressive compression strategy that exploits the higher compressibility of deeper-layer activations. Our method reduces computation cost, memory usage, and training time while maintaining accuracy. Experiments on NVIDIA Orin Edge GPU further demonstrate training acceleration and significant power savings, highlighting its practicality for resource-constrained training. Chence Yang, Ningxi Cheng, Ci Zhang, Qitao Tan, Sheng Li 0019, Ao Li 0004, Xulong Tang, Shaoyi Huang, Jinzhen Wang, Jundong Li, Xiaoming Zhai, Jin Lu 0001, Geng Yuan |
ACM Great Lakes Symposium on VLSI | 5 |
| 2025 | A Computation and Energy Efficient Hardware Architecture for SSL AccelerationabstractIn Computer Vision (CV), the deployment of Convolutional Neural Networks (CNNs) is often hindered by their substantial computational requirements and large labeled datasets. Self-supervised learning (SSL) serves as an effective approach to reducing the reliance on labeled data with the option of augmentation methods to infer and train CNNs. Excluding irrelevant features accelerates learning and improves optimization. We propose a Field-Programmable Gate Array (FPGA)-based hardware accelerator architecture tailored for SSL framework, leveraging its parallelism and reconfigurability to expedite block matching, optimize sparse convolutions, and manage data reuse, significantly improving resource and energy efficiency. The implementation and evaluation of our work on Xilinx ZCU102 FPGA working at 200 MHz confirm that the similarity finding part's FPGA accelerations with a low hardware overhead generates a latency of 0.0106 seconds, surpassing GPU and CPU, and in the sparse CNN's FPGA acceleration part, with the processing of VGG16 and ResNet50, compared with the related FPGA-based works, our design claims a maximum of 3.08× throughput improvement and 1.5× in energy efficiency. Huidong Ji, Sheng Li 0019, Chen Ding 0010, Jiawei Xu 0001, Qitao Tan, Jun Liu 0075, Ao Li 0004, Xulong Tang, Lirong Zheng 0001, Geng Yuan, Zhuo Zou |
ASP-DAC | 6 |
| 2025 | Towards Memory-Efficient and Sustainable Machine Unlearning on Edge using Zeroth-Order Optimizer
Ci Zhang, Chence Yang, Qitao Tan, Jun Liu 0075, Ao Li 0004, Yanzhi Wang 0001, Jin Lu 0001, Geng Yuan |
ACM Great Lakes Symposium on VLSI | 3 |
| 2025 | Perturbation-efficient Zeroth-order Optimization for Hardware-friendly On-device TrainingabstractZeroth-order (ZO) optimization is an emerging deep neural network (DNN) training paradigm that offers computational simplicity and memory savings. However, this seemingly promising approach faces a significant and long-ignored challenge. ZO requires generating a substantial number of Gaussian random numbers, which poses significant difficulties and even makes it infeasible for hardware platforms, such as FPGAs and ASICs. In this paper, we identify this critical issue, which arises from the mismatch between algorithm and hardware designers. To address this issue, we proposed PeZO, a perturbation-efficient ZO framework. Specifically, we design random number reuse strategies to significantly reduce the demand for random number generation and introduce a hardware-friendly adaptive scaling method to replace the costly Gaussian distribution with a uniform distribution. Our experiments show that PeZO reduces the required LUTs and FFs for random number generation by 48.6% and 12.7%, and saves at maximum 86% power consumption, all without compromising training performance, making ZO optimization feasible for on-device training. To the best of our knowledge, we are the first to explore the potential of on-device ZO optimization, providing valuable insights for future research. Qitao Tan, Sung-En Chang, Huidong Ji, Chence Yang, Ci Zhang, Jun Liu 0075, Zheng Zhan 0001, Zhenman Fang, Zhuo Zou, Yanzhi Wang 0001, Jin Lu 0001, Geng Yuan |
ICCAD | 1 |
| 2025 | Mutual Effort for Efficiency: A Similarity-based Token Pruning for Vision Transformers in Self-Supervised LearningabstractSelf-supervised learning (SSL) offers a compelling solution to the challenge of extensive labeled data requirements in traditional supervised learning.
With the proven success of Vision Transformers (ViTs) in supervised tasks, there is increasing interest in adapting them for SSL frameworks. However, the high computational demands of SSL pose substantial challenges, particularly on resource-limited platforms like edge devices, despite its ability to achieve high accuracy without labeled data.
Recent studies in supervised learning have shown that token pruning can reduce training costs by removing less informative tokens without compromising accuracy. However, SSL’s dual-branch encoders make traditional single-branch pruning strategies less effective, as they fail to account for the critical cross-branch similarity information, leading to reduced accuracy in SSL.
To this end, we introduce SimPrune, a novel token pruning strategy designed for ViTs in SSL. SimPrune leverages cross-branch similarity information to efficiently prune tokens, retaining essential semantic information across dual branches. Additionally, we incorporate a difficulty-aware pruning strategy to further enhance SimPrune's effectiveness.
Experimental results show that our proposed approach effectively reduces training computation while maintaining accuracy. Specifically, our approach offers 24\% savings in training costs compared to SSL baseline, without sacrificing accuracy. Sheng Li 0019, Qitao Tan, Yue Dai 0005, Zhenglun Kong, Jun Liu 0075, Ao Li 0004, Ninghao Liu 0001, Yufei Ding 0001, Xulong Tang, Geng Yuan |
ICLR | 2 |
| 2025 | Harmony in Divergence: Towards Fast, Accurate, and Memory-efficient Zeroth-order LLM Fine-tuningabstractLarge language models (LLMs) excel across various tasks, but standard first-order (FO) fine-tuning demands considerable memory, significantly limiting real-world deployment. Recently, zeroth-order (ZO) optimization stood out as a promising memory-efficient training paradigm, avoiding backward passes and relying solely on forward passes for gradient estimation, making it attractive for resource-constrained scenarios. However, ZO method lags far behind FO method in both convergence speed and accuracy. To bridge the gap, we introduce a novel layer-wise divergence analysis that uncovers the distinct update pattern of FO and ZO optimization. Aiming to resemble the learning capacity of FO method from the findings, we propose \textbf{Di}vergence-driven \textbf{Z}eroth-\textbf{O}rder (\textbf{DiZO}) optimization. DiZO conducts divergence-driven layer adaptation by incorporating projections to ZO updates, generating diverse-magnitude updates precisely scaled to layer-wise individual optimization needs. Our results demonstrate that DiZO significantly reduces the needed iterations for convergence without sacrificing throughput, cutting training GPU hours by up to 48\% on various datasets. Moreover, DiZO consistently outperforms the representative ZO baselines in fine-tuning RoBERTa-large, OPT-series, and Llama-series on downstream tasks and, in some cases, even surpasses memory-intensive FO fine-tuning. Our code is released at \url{https://github.com/Skilteee/DiZO}. Qitao Tan, Jun Liu 0075, Zheng Zhan 0001, Caiwen Ding, Yanzhi Wang 0001, Jin Lu 0001, Geng Yuan |
NeurIPS | 1 |
| 2023 | An effective negative sampling approach for contrastive learning of sentence embedding
Qitao Tan, Guanghui Ye, Chuan Wu 0003 |
Mach. Learn. | 1 |