EDBT 2026 Demo / reviewers in the wild / expert
Yuhang Li 0001
dblp:72/2508-1
· DBLP profile ↗
33ranked-venue papers
13as first author
31since 2021 · last 2025
0000-0002-6444-7253ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 12 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 12 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Spiking Transformer with Spatial-Temporal AttentionabstractSpike-based Transformer presents a compelling and energy-efficient alternative to traditional Artificial Neural Network (ANN)-based Transformers, achieving impressive results through sparse binary computations. However, existing spike-based transformers predominantly focus on spatial attention while neglecting crucial temporal dependencies inherent in spike-based processing, leading to suboptimal feature representation and limited performance. To address this limitation, we propose Spiking Transformer with Spatial-Temporal Attention (STAtten), a simple and straightforward architecture that efficiently integrates both spatial and temporal information in the self-attention mechanism. STAtten introduces a block-wise computation strategy that processes information in spatial-temporal chunks, enabling comprehensive feature capture while maintaining the same computational complexity as previous spatial-only approaches. Our method can be seamlessly integrated into existing spike-based transformers without architectural overhaul. Extensive experiments demonstrate that STAtten significantly improves the performance of existing spike-based transformers across both static and neuromorphic datasets, including CIFAR10/100, ImageNet, CIFAR10-DVS, and N-Caltech101. The code is available at https://github.com/Intelligent-Computing-Lab-Yale/STAtten. Donghyun Lee 0002, Yuhang Li 0001, Youngeun Kim, Shiting Xiao, Priyadarshini Panda |
CVPR | 2 |
| 2025 | PacQ: A SIMT Microarchitecture for Efficient Dataflow in Hyper-asymmetric GEMMsabstractWeight-only quantization has been widely explored in large language models (LLMs) to reduce memory storage and data loading overhead. During deployment on single-instruction-multiple-threads (SIMT) architectures, weights are stored in low-precision integer (INT) format, while activations remain in full-precision floating-point (FP) format to preserve inference accuracy. Although memory footprint and data loading requirements for weight matrices are reduced, computation performance gains remain limited due to the need to convert weights back to FP format through unpacking and dequantization before GEMM operations. In this work, we investigate methods to accelerate GEMM operations involving packed low-precision INT weights and high-precision FP activations, defining this as the hyper-asymmetric GEMM problem. Our approach co-optimizes tile-level packing and dataflow strategies for INT weight matrices. We further design a specialized FP-INT multiplier unit tailored to our packing and dataflow strategies, enabling parallel processing of multiple INT weights. Finally, we integrate the packing, dataflow, and multiplier unit into PacQ, a SIMT microarchitecture designed to efficiently accelerate hyper-asymmetric GEMMs. We show that PacQ can achieve up to $1.99 \times$ speedup and $81.4 \%$ reduction in EDP compared to weight-only quantized LLM workloads running on conventional SIMT baselines. Ruokai Yin, Yuhang Li 0001, Priyadarshini Panda |
DAC | 2 |
| 2025 | GPTAQ: Efficient Finetuning-Free Quantization for Asymmetric CalibrationabstractWe introduce GPTAQ, a novel finetuning-free quantization method for compressing large-scale transformer architectures.
Unlike the previous GPTQ method, which independently calibrates each layer, we always match the quantized layer's output to the exact output in the full-precision model, resulting in a scheme that we call *asymmetric calibration*. Such a scheme can effectively reduce the quantization error accumulated in previous layers. We analyze this problem using optimal brain compression to derive a close-formed solution. The new solution explicitly minimizes the quantization error as well as the accumulated asymmetry error. Furthermore, we utilize various techniques to parallelize the solution calculation, including channel parallelization, neuron decomposition, and Cholesky reformulation for matrix fusion. As a result, GPTAQ is easy to implement, simply using 20 more lines of code than GPTQ but improving its performance under low-bit quantization. Remarkably, on a single GPU, we quantize a 405B language transformer as well as EVA-02—the rank first vision transformer that achieves 90% pretraining Imagenet accuracy. Code is available at [Github](https://github.com/Intelligent-Computing-Lab-Yale/GPTAQ). Yuhang Li 0001, Ruokai Yin, Donghyun Lee 0002, Shiting Xiao, Priyadarshini Panda |
ICML | 1 |
| 2025 | OpenWorldSAM: Extending SAM2 for Universal Image Segmentation with Language PromptsabstractThe ability to segment objects based on open-ended language prompts remains a critical challenge, requiring models to ground textual semantics into precise spatial masks while handling diverse and unseen categories. We present OpenWorldSAM, a framework that extends the prompt-driven Segment Anything Model v2 (SAM2) to open-vocabulary scenarios by integrating multi-modal embeddings extracted from a lightweight vision-language model (VLM). Our approach is guided by four key principles: i) Unified prompting: OpenWorldSAM supports a diverse range of prompts, including category-level and sentence-level language descriptions, providing a flexible interface for various segmentation tasks. ii) Efficiency: By freezing the pre-trained components of SAM2 and the VLM, we train only 4.5 million parameters on the COCO-stuff dataset, achieving remarkable resource efficiency. iii) Instance Awareness: We enhance the model's spatial understanding through novel positional tie-breaker embeddings and cross-attention layers, enabling effective segmentation of multiple instances. iv) Generalization: OpenWorldSAM exhibits strong zero-shot capabilities, generalizing well on unseen categories and an open vocabulary of concepts without additional training. Extensive experiments demonstrate that OpenWorldSAM achieves state-of-the-art performance in open-vocabulary semantic, instance, and panoptic segmentation across multiple benchmarks. Code is available at https://github.com/GinnyXiao/OpenWorldSAM. Shiting Xiao, Rishabh Kabra, Yuhang Li 0001, Donghyun Lee 0002, João Carreira 0001, Priyadarshini Panda |
NeurIPS | 3 |
| 2025 | DuoGPT: Training-free Dual Sparsity through Activation-aware Pruning in LLMsabstractLarge language models (LLMs) deliver strong performance but are difficult to deploy due to high memory and compute costs. While pruning reduces these demands, most methods ignore activation sparsity observed at runtime. We reinterpret activation sparsity as dynamic structured weight sparsity and propose DuoGPT, a unified framework that constructs dual-sparse (spMspV) workloads by combining unstructured weight pruning with activation sparsity. To preserve accuracy, we extend the Optimal Brain Compression (OBC) framework with activation-aware calibration and introduce output residuals from the dense model as correction terms. We further optimize the solution for efficient GPU execution, enabling scalability to billion-parameter LLMs. Evaluations on LLaMA-2 and LLaMA-3 show that DuoGPT outperforms state-of-the-art structured pruning methods by up to 9.17\% accuracy at an iso-speedup of 1.39$\times$ compared to the baseline dense model. Code is available at GitHub. Ruokai Yin, Yuhang Li 0001, Donghyun Lee 0002, Priyadarshini Panda |
NeurIPS | 2 |
| 2025 | Pushing the Limit of Post-Training QuantizationabstractRecently, post-training quantization (PTQ) has become the de facto way to produce efficient low-precision neural networks without long-time retraining. Despite its low cost, current PTQ works fail to succeed under the extremely low-bit setting. In this work, we delve into extremely low-bit quantization and construct a unified theoretical analysis, which provides an in-depth understanding of the reason for the failure of low-bit quantization. According to the theoretical study, we argue that the existing methods fail in low-bit schemes due to significant perturbation on weights and lack of consideration of activation quantization. To this end, we propose Brecq and QDrop to respectively solve these two challenges, based on which a Q-Limit framework is constructed. Then the Q-Limit framework is further extended to support a mixed precision quantization scheme. To the best of our knowledge, this is the first work that can push the limit of PTQ down to INT2. Extensive experiments on various handcrafted and searched neural architectures are conducted for both visual recognition/detection tasks and language processing tasks. Without bells and whistles, our PTQ framework can attain low-bit ResNet and MobileNetV2 comparable with quantization-aware training (QAT), establishing a new state-of-the-art for PTQ. Ruihao Gong, Xianglong Liu 0001, Yuhang Li 0001, Yunqian Fan, Xiuying Wei, Jinyang Guo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Temporal Feature Matters: A Framework for Diffusion Model QuantizationabstractDiffusion models, widely used for image generation, face significant challenges related to their broad applicability due to prolonged inference times and high memory demands. Efficient Post-Training Quantization (PTQ) is crucial to address these issues. However, unlike traditional models, diffusion models critically rely on the time-step for the multi-round denoising. Typically, each time-step is encoded into a hypersensitive temporal feature by several modules. Despite this, existing PTQ methods do not optimize these modules individually. Instead, they employ unsuitable reconstruction objectives and complex calibration methods, leading to significant disturbances in the temporal feature and denoising trajectory, as well as reduced compression efficiency. To address these challenges, we introduce a novel quantization framework that includes three strategies: 1) TIB-based Maintenance: Based on our innovative Temporal Information Block (TIB) definition, Temporal Information-aware Reconstruction (TIAR) and Finite Set Calibration (FSC) are developed to efficiently align original temporal features. 2) Cache-based Maintenance: Instead of indirect and complex optimization for the related modules, pre-computing and caching quantized counterparts of temporal features are developed to minimize errors. 3) Disturbance-aware Selection: Employ temporal feature errors to guide a fine-grained selection between the two maintenance strategies for further disturbance reduction. This framework preserves most of the temporal information and ensures high-quality end-to-end generation. Extensive testing on various datasets, diffusion models and hardware confirms our superior performance and acceleration. Yushi Huang, Ruihao Gong, Xianglong Liu 0001, Jing Liu 0048, Yuhang Li 0001, Jiwen Lu, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | MINT: Multiplier-less INTeger Quantization for Energy Efficient Spiking Neural NetworksabstractWe propose Multiplier-less INTeger (MINT) quantization, a uniform quantization scheme that efficiently compresses weights and membrane potentials in spiking neural networks (SNNs). Unlike previous SNN quantization methods, MINT quantizes memory-intensive membrane potentials to an extremely low precision (2-bit), significantly reducing the memory footprint. MINT also shares the quantization scaling factor between weights and membrane potentials, eliminating the need for multipliers required in conventional uniform quantization. Experimental results show that our method matches the accuracy of full-precision models and other state-of-the-art SNN quantization techniques while surpassing them in memory footprint reduction and hardware cost efficiency at deployment. For example, 2-bit MINT VGG-16 achieves 90.6% accuracy on CIFAR-10, with roughly 93.8% reduction in memory footprint from the full-precision model and 90% reduction in computation energy compared to vanilla uniform quantization at deployment.11Code is available at https://github.com/Intelligent-Computing-Lab-Yale/MINT-Quantization Ruokai Yin, Yuhang Li 0001, Abhishek Moitra, Priyadarshini Panda |
ASPDAC | 2 |
| 2024 | TT-SNN: Tensor Train Decomposition for Efficient Spiking Neural Network TrainingabstractSpiking Neural Networks (SNNs) have gained significant attention as a potentially energy-efficient alternative for standard neural networks with their sparse binary activation. However, SNNs suffer from memory and computation overhead due to spatio-temporal dynamics and multiple backpropagation computations across timesteps during training. To address this issue, we introduce Tensor Train Decomposition for Spiking Neural Networks (TT-SNN), a method that reduces model size through trainable weight decomposition, resulting in reduced storage, FLOPs, and latency. In addition, we propose a parallel computation pipeline as an alternative to the typical sequential tensor computation, which can be flexibly integrated into various existing SNN architectures. To the best of our knowledge, this is the first of its kind application of tensor decomposition in SNNs. We validate our method using both static and dynamic datasets, CIFAR1I0/100 and N-Caltechl0l, respectively. We also propose a TT-SNN-tailored training accelerator to fully harness the parallelism in TT-SNN. Our results demonstrate substantial reductions in parameter size$(7.98\times)$, FLOPs$(9.25\times)$, training time (17.7 %), and training energy (28.3 %) during training for the N-Caltechl0l dataset, with negligible accuracy degradation. Donghyun Lee 0002, Ruokai Yin, Youngeun Kim, Abhishek Moitra, Yuhang Li 0001, Priyadarshini Panda |
DATE | 5 |
| 2024 | One-Stage Prompt-Based Continual Learning
Youngeun Kim, Yuhang Li 0001, Priyadarshini Panda |
ECCV (13) | 2 |
| 2024 | A Simple Background Augmentation Method for Object Detection with Diffusion Model
Yuhang Li 0001, Xin Dong 0009, Chen Chen 0043, Weiming Zhuang, Lingjuan Lyu |
ECCV (66) | 1 |
| 2024 | GenQ: Quantization in Low Data Regimes with Generative Synthetic Data
Yuhang Li 0001, Youngeun Kim, Donghyun Lee 0002, Souvik Kundu 0009, Priyadarshini Panda |
ECCV (13) | 1 |
| 2024 | Error-Aware Conversion from ANN to SNN via Post-training Parameter Calibration
Yuhang Li 0001, Shikuang Deng, Xin Dong 0009, Shi Gu |
Int. J. Comput. Vis. | 1 |
| 2024 | Do we really need a large number of visual prompts?
Youngeun Kim, Yuhang Li 0001, Abhishek Moitra, Ruokai Yin, Priyadarshini Panda |
Neural Networks | 2 |
| 2023 | Exploring Temporal Information Dynamics in Spiking Neural NetworksabstractMost existing Spiking Neural Network (SNN) works state that SNNs may utilize temporal information dynamics of spikes. However, an explicit analysis of temporal information dynamics is still missing. In this paper, we ask several important questions for providing a fundamental understanding of SNNs: What are temporal information dynamics inside SNNs? How can we measure the temporal information dynamics? How do the temporal information dynamics affect the overall learning performance? To answer these questions, we estimate the Fisher Information of the weights to measure the distribution of temporal information during training in an empirical manner. Surprisingly, as training goes on, Fisher information starts to concentrate in the early timesteps. After training, we observe that information becomes highly concentrated in earlier few timesteps, a phenomenon we refer to as temporal information concentration. We observe that the temporal information concentration phenomenon is a common learning feature of SNNs by conducting extensive experiments on various configurations such as architecture, dataset, optimization strategy, time constant, and timesteps. Furthermore, to reveal how temporal information concentration affects the performance of SNNs, we design a loss function to change the trend of temporal information. We find that temporal information concentration is crucial to building a robust SNN but has little effect on classification accuracy. Finally, we propose an efficient iterative pruning method based on our observation on temporal information concentration. Code is available at https://github.com/Intelligent-Computing-Lab-Yale/Exploring-Temporal-Information-Dynamics-in-Spiking-Neural-Networks. Youngeun Kim, Yuhang Li 0001, Hyoungseob Park, Yeshwanth Venkatesha, Anna Hambitzer, Priyadarshini Panda |
AAAI | 2 |
| 2023 | Input-Aware Dynamic Timestep Spiking Neural Networks for Efficient In-Memory ComputingabstractSpiking Neural Networks (SNNs) have recently attracted widespread research interest as an efficient alternative to traditional Artificial Neural Networks (ANNs) because of their capability to process sparse and binary spike information and avoid expensive multiplication operations. Although the efficiency of SNNs can be realized on the In-Memory Computing (IMC) architecture, we show that the energy cost and latency of SNNs scale linearly with the number of timesteps used on IMC hardware. Therefore, in order to maximize the efficiency of SNNs, we propose input-aware Dynamic Timestep SNN (DT-SNN), a novel algorithmic solution to dynamically determine the number of timesteps during inference on an input-dependent basis. By calculating the entropy of the accumulated output after each timestep, we can compare it to a predefined threshold and decide if the information processed at the current timestep is sufficient for a confident prediction. We deploy DT-SNN on an IMC architecture and show that it incurs negligible computational overhead. We demonstrate that our method only uses 1.46 average timesteps to achieve the accuracy of a 4-timestep static SNN while reducing the energy-delay-product by 80%. Yuhang Li 0001, Abhishek Moitra, Tamar Geller, Priyadarshini Panda |
DAC | 1 |
| 2023 | Outlier Suppression+: Accurate quantization of large language models by equivalent and effective shifting and scalingabstractPost-training quantization (PTQ) of transformer language models faces significant challenges due to the existence of detrimental outliers in activations.We observe that these outliers are concentrated in specific channels and are asymmetric across channels.To address this issue, we propose the Outlier Suppression+ (OS+) framework, which contains the channel-wise shifting for asymmetry and channel-wise scaling for concentration.We show that these operations can be seamlessly migrated into subsequent modules while maintaining equivalence.Second, we propose a fast and stable scheme to calculate effective shifting and scaling values.The channel-wise shifting aligns the center of each channel for removal of outlier asymmetry.The channel-wise scaling quantitatively evaluates changes brought by migration and quantization for better quantization burden balance.We validate our OS+ under both standard and fine-grained quantization settings with models including BERT, OPT, BLOOM, BLOOMZ, and LLaMA.Comprehensive results across various tasks demonstrate the superiority of our approach.Especially, with standard quantization, OS+ can achieve near-floating-point performance on both small models and large language models on 8-bit and 6-bit.Besides, we establish a new state-of-the-art for 4-bit BERT with 15.5% improvement.Our code is available at https://github.com/ModelTC/ Outlier_Suppression_Plus. Xiuying Wei, Yunchen Zhang, Yuhang Li 0001, Xiangguo Zhang, Ruihao Gong, Jinyang Guo 0002, Xianglong Liu 0001 |
EMNLP | 3 |
| 2023 | Augmentation Robust Self-Supervised Learning for Human Activity RecognitionabstractHuman Activity Recognition (HAR) is widely applied on wearable devices in our daily lives. However, acquiring high-quality wearable sensor data set with ground-truths is challenging due to the high cost in collecting data and necessity of domain experts. In order to achieve generalization from limited data, we study augmentation-based Self-Supervised Learning (SSL) for data from wearable devices. However, there is an issue in one of the most popular SSL approaches, contrastive learning: it is sensitive to the choice of data augmentations. To resolve this, we first propose to combine contrastive learning with generative learning, which is robust to augmentations. Second, we propose an automatic augmentation policy search method to discover the most promising augmentation policy. We empirically verify our approaches on three public HAR datasets. Experimental results show that our proposed SSL approach is robust to augmentations, and delivers higher accuracy than contrastive learning. Additionally, with the searched augmentation policy we are able to further improve the accuracy of HAR task. Yuhang Li 0001, Dae Lee, Dae Hoon Park, Hongda Mao, Huyen Do, Jonathan Chung 0001, Dinesh Nair |
ICASSP | 2 |
| 2023 | Surrogate Module Learning: Reduce the Gradient Error Accumulation in Training Spiking Neural NetworksabstractSpiking neural networks provide an alternative solution to conventional artificial neural networks with energy-saving and high-efficiency characteristics after hardware implantation. However, due to its non-differentiable activation function and the temporally delayed accumulation in outputs, the direct training of SNNs is extraordinarily tough even adopting a surrogate gradient to mimic the backpropagation. For SNN training, this non-differentiability causes the intrinsic gradient error that would be magnified through layerwise backpropagation, especially through multiple layers. In this paper, we propose a novel approach to reducing gradient error from a new perspective called surrogate module learning (SML). Surrogate module learning tries to construct a shortcut path to back-propagate more accurate gradient to a certain SNN part utilizing the surrogate modules. Then, we develop a new loss function for concurrently training the network and enhancing the surrogate modules' surrogate capacity. We demonstrate that when the outputs of surrogate modules are close to the SNN output, the fraction of the gradient error drops significantly. Our method consistently and significantly enhances the performance of SNNs on all experiment datasets, including CIFAR-10/100, ImageNet, and ES-ImageNet. For example, for spiking ResNet-34 architecture on ImageNet, we increased the SNN accuracy by 3.46%. Shikuang Deng, Yuhang Li 0001, Shi Gu |
ICML | 3 |
| 2023 | SEENN: Towards Temporal Spiking Early Exit Neural NetworksabstractSpiking Neural Networks (SNNs) have recently become more popular as a biologically plausible substitute for traditional Artificial Neural Networks (ANNs). SNNs are cost-efficient and deployment-friendly because they process input in both spatial and temporal manner using binary spikes. However, we observe that the information capacity in SNNs is affected by the number of timesteps, leading to an accuracy-efficiency tradeoff. In this work, we study a fine-grained adjustment of the number of timesteps in SNNs. Specifically, we treat the number of timesteps as a variable conditioned on different input samples to reduce redundant timesteps for certain data.
We call our method Spiking Early-Exit Neural Networks (**SEENNs**). To determine the appropriate number of timesteps, we propose SEENN-I which uses a confidence score thresholding to filter out the uncertain predictions, and SEENN-II which determines the number of timesteps by reinforcement learning.
Moreover, we demonstrate that SEENN is compatible with both the directly trained SNN and the ANN-SNN conversion.
By dynamically adjusting the number of timesteps, our SEENN achieves a remarkable reduction in the average number of timesteps during inference. For example, our SEENN-II ResNet-19 can achieve **96.1**\% accuracy with an average of **1.08** timesteps on the CIFAR-10 test dataset. Code is shared at https://github.com/Intelligent-Computing-Lab-Yale/SEENN. Yuhang Li 0001, Tamar Geller, Youngeun Kim, Priyadarshini Panda |
NeurIPS | 1 |
| 2022 | Neural Architecture Search for Spiking Neural Networks
Youngeun Kim, Yuhang Li 0001, Hyoungseob Park, Yeshwanth Venkatesha, Priyadarshini Panda |
ECCV (24) | 2 |
| 2022 | Exploring Lottery Ticket Hypothesis in Spiking Neural Networks
Youngeun Kim, Yuhang Li 0001, Hyoungseob Park, Yeshwanth Venkatesha, Ruokai Yin, Priyadarshini Panda |
ECCV (12) | 2 |
| 2022 | Neuromorphic Data Augmentation for Training Spiking Neural Networks
Yuhang Li 0001, Youngeun Kim, Hyoungseob Park, Tamar Geller, Priyadarshini Panda |
ECCV (7) | 1 |
| 2022 | Temporal Efficient Training of Spiking Neural Network via Gradient Re-weighting
Shikuang Deng, Yuhang Li 0001, Shanghang Zhang, Shi Gu |
ICLR | 2 |
| 2022 | QDrop: Randomly Dropping Quantization for Extremely Low-bit Post-Training Quantization
Xiuying Wei, Ruihao Gong, Yuhang Li 0001, Xianglong Liu 0001, Fengwei Yu |
ICLR | 3 |
| 2021 | Diversifying Sample Generation for Accurate Data-Free QuantizationabstractQuantization has emerged as one of the most prevalent approaches to compress and accelerate neural networks. Recently, data-free quantization has been widely studied as a practical and promising solution. It synthesizes data for calibrating the quantized model according to the batch normalization (BN) statistics of FP32 ones and significantly relieves the heavy dependency on real training data in traditional quantization methods. Unfortunately, we find that in practice, the synthetic data identically constrained by BN statistics suffers serious homogenization at both distribution level and sample level and further causes a significant performance drop of the quantized model. We propose Diverse Sample Generation (DSG) scheme to mitigate the adverse effects caused by homogenization. Specifically, we slack the alignment of feature statistics in the BN layer to relax the constraint at the distribution level and design a layerwise enhancement to reinforce specific layers for different data samples. Our DSG scheme is versatile and even able to be applied to the state-of-the-art post-training quantization method like AdaRound. We evaluate the DSG scheme on the large-scale image classification task and consistently obtain significant improvements over various network architectures and quantization methods, especially when quantized to lower bits (e.g., up to 22% improvement on W4A4). Moreover, benefiting from the enhanced diversity, models calibrated with synthetic data perform close to those calibrated with real data and even outperform them on W4A4. Xiangguo Zhang, Haotong Qin, Yifu Ding 0001, Ruihao Gong, Qinghua Yan, Renshuai Tao, Yuhang Li 0001, Fengwei Yu, Xianglong Liu 0001 |
CVPR | 7 |
| 2021 | MixMix: All You Need for Data-Free Compression Are Feature and Data MixingabstractUser data confidentiality protection is becoming a rising challenge in the present deep learning research. Without access to data, conventional data-driven model compression faces a higher risk of performance degradation. Recently, some works propose to generate images from a specific pretrained model to serve as training data. However, the inversion process only utilizes biased feature statistics stored in one model and is from low-dimension to high-dimension. As a consequence, it inevitably encounters the difficulties of generalizability and inexact inversion, which leads to unsatisfactory performance. To address these problems, we propose MixMix based on two simple yet effective techniques: (1) Feature Mixing: utilizes various models to construct a universal feature space for generalized inversion; (2) Data Mixing: mixes the synthesized images and labels to generate exact label information. We prove the effectiveness of MixMix from both theoretical and empirical perspectives. Extensive experiments show that MixMix outperforms existing methods on the mainstream compression tasks, including quantization, knowledge distillation and pruning. Specifically, MixMix achieves up to 4% and 20% accuracy uplift on quantization and pruning, respectively, compared to existing data-free compression work. Yuhang Li 0001, Feng Zhu 0006, Ruihao Gong, Mingzhu Shen, Xin Dong 0009, Fengwei Yu, Shaoqing Lu, Shi Gu |
ICCV | 1 |
| 2021 | Once Quantization-Aware Training: High Performance Extremely Low-bit Architecture SearchabstractQuantization Neural Networks (QNN) have attracted a lot of attention due to their high efficiency. To enhance the quantization accuracy, prior works mainly focus on designing advanced quantization algorithms but still fail to achieve satisfactory results under the extremely low-bit case. In this work, we take an architecture perspective to investigate the potential of high-performance QNN. Therefore, we propose to combine Network Architecture Search methods with quantization to enjoy the merits of the two sides. However, a naive combination inevitably faces unacceptable time consumption or unstable training problem. To alleviate these problems, we first propose the joint training of architecture and quantization with a shared step size to acquire a large number of quantized models. Then a bit-inheritance scheme is introduced to transfer the quantized models to the lower bit, which further reduces the time cost and meanwhile improves the quantization accuracy. Equipped with this overall framework, dubbed as Once Quantization-Aware Training (OQAT), our searched model family, OQATNets, achieves a new state-of-the-art compared with various architectures under different bit-widths. In particular, OQAT-2bit-M achieves 61.6% ImageNet Top-1 accuracy, outperforming 2-bit counterpart MobileNetV3 by a large margin of 9% with 10% less computation cost. A series of quantization-friendly architectures are identified easily and extensive analysis can be made to summarize the interaction between quantization and neural architectures. Codes and models are released at https://github.com/LaVieEnRoseSMZ/OQA Mingzhu Shen, Ruihao Gong, Yuhang Li 0001, Chuming Li, Chen Lin 0003, Fengwei Yu, Wanli Ouyang |
ICCV | 4 |
| 2021 | BRECQ: Pushing the Limit of Post-Training Quantization by Block Reconstruction
Yuhang Li 0001, Ruihao Gong, Fengwei Yu, Wei Wang 0059, Shi Gu |
ICLR | 1 |
| 2021 | A Free Lunch From ANN: Towards Efficient, Accurate Spiking Neural Networks CalibrationabstractSpiking Neural Network (SNN) has been recognized as one of the next generation of neural networks. Conventionally, SNN can be converted from a pre-trained ANN by only replacing the ReLU activation to spike activation while keeping the parameters intact. Perhaps surprisingly, in this work we show that a proper way to calibrate the parameters during the conversion of ANN to SNN can bring significant improvements. We introduce SNN Calibration, a cheap but extraordinarily effective method by leveraging the knowledge within a pre-trained Artificial Neural Network (ANN). Starting by analyzing the conversion error and its propagation through layers theoretically, we propose the calibration algorithm that can correct the error layer-by-layer. The calibration only takes a handful number of training data and several minutes to finish. Moreover, our calibration algorithm can produce SNN with state-of-the-art architecture on the large-scale ImageNet dataset, including MobileNet and RegNet. Extensive experiments demonstrate the effectiveness and efficiency of our algorithm. For example, our advanced pipeline can increase up to 69% top-1 accuracy when converting MobileNet on ImageNet compared to baselines. Codes are released at https://github.com/yhhhli/SNN_Calibration. Yuhang Li 0001, Shikuang Deng, Xin Dong 0009, Ruihao Gong, Shi Gu |
ICML | 1 |
| 2021 | Differentiable Spike: Rethinking Gradient-Descent for Training Spiking Neural NetworksabstractSpiking Neural Networks (SNNs) have emerged as a biology-inspired method mimicking the spiking nature of brain neurons. This bio-mimicry derives SNNs' energy efficiency of inference on neuromorphic hardware. However, it also causes an intrinsic disadvantage in training high-performing SNNs from scratch since the discrete spike prohibits the gradient calculation. To overcome this issue, the surrogate gradient (SG) approach has been proposed as a continuous relaxation. Yet the heuristic choice of SG leaves it vacant how the SG benefits the SNN training. In this work, we first theoretically study the gradient descent problem in SNN training and introduce finite difference gradient to quantitatively analyze the training behavior of SNN. Based on the introduced finite difference gradient, we propose a new family of Differentiable Spike (Dspike) functions that can adaptively evolve during training to find the optimal shape and smoothness for gradient estimation. Extensive experiments over several popular network structures show that training SNN with Dspike consistently outperforms the state-of-the-art training methods. For example, on the CIFAR10-DVS classification task, we can train a spiking ResNet-18 and achieve 75.4% top-1 accuracy with 10 time steps. Yuhang Li 0001, Yufei Guo 0001, Shanghang Zhang, Shikuang Deng, Yongqing Hai, Shi Gu |
NeurIPS | 1 |
| 2020 | RTN: Reparameterized Ternary NetworkabstractTo deploy deep neural networks on resource-limited devices, quantization has been widely explored. In this work, we study the extremely low-bit networks which have tremendous speed-up, memory saving with quantized activation and weights. We first bring up three omitted issues in extremely low-bit networks: the squashing range of quantized values; the gradient vanishing during backpropagation and the unexploited hardware acceleration of ternary networks. By reparameterizing quantized activation and weights vector with full precision scale and offset for fixed ternary vector, we decouple the range and magnitude from direction to extenuate above problems. Learnable scale and offset can automatically adjust the range of quantized values and sparsity without gradient vanishing. A novel encoding and computation pattern are designed to support efficient computing for our reparameterized ternary network (RTN). Experiments on ResNet-18 for ImageNet demonstrate that the proposed RTN finds a much better efficiency between bitwidth and accuracy and achieves up to 26.76% relative accuracy improvement compared with state-of-the-art methods. Moreover, we validate the proposed computation pattern on Field Programmable Gate Arrays (FPGA), and it brings 46.46 × and 89.17 × savings on power and area compared with the full precision convolution. Yuhang Li 0001, Xin Dong 0009, Sai Qian Zhang, Haoli Bai, Yuanpeng Chen, Wei Wang 0059 |
AAAI | 1 |
| 2020 | Additive Powers-of-Two Quantization: An Efficient Non-uniform Discretization for Neural Networks
Yuhang Li 0001, Xin Dong 0009, Wei Wang 0059 |
ICLR | 1 |