EDBT 2026 Demo / reviewers in the wild / expert
Ruokai Yin
dblp:270/3836
· DBLP profile ↗
15ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0002-7550-0638ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MD-SNN: Membrane Potential-aware Distillation on Quantized Spiking Neural NetworkabstractSpiking Neural Networks (SNNs) offer a promising and energy-efficient alternative to conventional neural networks, thanks to their sparse binary activation. However, they face challenges regarding memory and computation overhead due to complex spatio-temporal dynamics and the necessity for multiple backpropagation computations across timesteps during training. To mitigate this overhead, compression techniques such as quantization are applied to SNNs. Yet, naively applying quantization to SNNs introduces a mismatch in membrane potential, a crucial factor for the firing of spikes, resulting in accuracy degradation. In this paper, we introduce Membrane-aware Distillation on quantized Spiking Neural Network (MD-SNN), which leverages membrane potential to mitigate discrepancies after weight, membrane potential, and batch normalization quantization. To our knowledge, this study represents the first application of membrane potential knowledge distillation in SNNs. We validate our approach on various datasets, including CIFAR10, CIFAR100, N-Caltech101, and TinyImageNet, demonstrating its effectiveness for both static and dynamic data scenarios. Furthermore, for hardware efficiency, we evaluate the MD-SNN with SpikeSim platform, finding that MD-SNNs achieve 14.85× lower energy-delay-area product (EDAP), 2.64× higher TOPS/W, and 6.19× higher TOPS/mm2compared to floating point SNNs at iso-accuracy on N-Caltech101 dataset. Code is available at Github. Donghyun Lee 0002, Abhishek Moitra, Youngeun Kim, Ruokai Yin, Priyadarshini Panda |
DATE | 4 |
| 2025 | PacQ: A SIMT Microarchitecture for Efficient Dataflow in Hyper-asymmetric GEMMsabstractWeight-only quantization has been widely explored in large language models (LLMs) to reduce memory storage and data loading overhead. During deployment on single-instruction-multiple-threads (SIMT) architectures, weights are stored in low-precision integer (INT) format, while activations remain in full-precision floating-point (FP) format to preserve inference accuracy. Although memory footprint and data loading requirements for weight matrices are reduced, computation performance gains remain limited due to the need to convert weights back to FP format through unpacking and dequantization before GEMM operations. In this work, we investigate methods to accelerate GEMM operations involving packed low-precision INT weights and high-precision FP activations, defining this as the hyper-asymmetric GEMM problem. Our approach co-optimizes tile-level packing and dataflow strategies for INT weight matrices. We further design a specialized FP-INT multiplier unit tailored to our packing and dataflow strategies, enabling parallel processing of multiple INT weights. Finally, we integrate the packing, dataflow, and multiplier unit into PacQ, a SIMT microarchitecture designed to efficiently accelerate hyper-asymmetric GEMMs. We show that PacQ can achieve up to $1.99 \times$ speedup and $81.4 \%$ reduction in EDP compared to weight-only quantized LLM workloads running on conventional SIMT baselines. Ruokai Yin, Yuhang Li 0001, Priyadarshini Panda |
DAC | 1 |
| 2025 | GPTAQ: Efficient Finetuning-Free Quantization for Asymmetric CalibrationabstractWe introduce GPTAQ, a novel finetuning-free quantization method for compressing large-scale transformer architectures.
Unlike the previous GPTQ method, which independently calibrates each layer, we always match the quantized layer's output to the exact output in the full-precision model, resulting in a scheme that we call *asymmetric calibration*. Such a scheme can effectively reduce the quantization error accumulated in previous layers. We analyze this problem using optimal brain compression to derive a close-formed solution. The new solution explicitly minimizes the quantization error as well as the accumulated asymmetry error. Furthermore, we utilize various techniques to parallelize the solution calculation, including channel parallelization, neuron decomposition, and Cholesky reformulation for matrix fusion. As a result, GPTAQ is easy to implement, simply using 20 more lines of code than GPTQ but improving its performance under low-bit quantization. Remarkably, on a single GPU, we quantize a 405B language transformer as well as EVA-02—the rank first vision transformer that achieves 90% pretraining Imagenet accuracy. Code is available at [Github](https://github.com/Intelligent-Computing-Lab-Yale/GPTAQ). Yuhang Li 0001, Ruokai Yin, Donghyun Lee 0002, Shiting Xiao, Priyadarshini Panda |
ICML | 2 |
| 2025 | SITRA: Exploiting Temporal Silence in Spiking Transformers for Fast & Energy-efficient InferenceabstractSpiking Neural Network (SNN) transformers are emerging as a compelling alternative to conventional Artificial Neural Network (ANN) transformers, offering enhanced energy-efficiency through binary spike-based computations and high temporal sparsity. However, a key limitation of SNN transformers is their reliance on sequential processing across multiple timesteps, resulting in a latency drawback. This work introduces SITRA, a hardware evaluation engine tailored for SNN transformers, which employs timestep-parallelization to eliminate the latency bottleneck. Furthermore, SITRA dynamically exploits the inherent sparsity of silent inputs and preemptively detects silent outputs during SNN transformer inference. This enables skipping unnecessary costly memory accesses and operations, thereby achieving significant energy-efficiency and latency improvements over ANN transformers. Using SITRA, we find upto 6.4× higher energy-efficiency and 3.4× lower inference latency when benchmarking a SpikingBERT language transformer against its ANN baseline. Abhiroop Bhattacharjee, Abhishek Moitra, Ruokai Yin, Priyadarshini Panda |
ISLPED | 3 |
| 2025 | DuoGPT: Training-free Dual Sparsity through Activation-aware Pruning in LLMsabstractLarge language models (LLMs) deliver strong performance but are difficult to deploy due to high memory and compute costs. While pruning reduces these demands, most methods ignore activation sparsity observed at runtime. We reinterpret activation sparsity as dynamic structured weight sparsity and propose DuoGPT, a unified framework that constructs dual-sparse (spMspV) workloads by combining unstructured weight pruning with activation sparsity. To preserve accuracy, we extend the Optimal Brain Compression (OBC) framework with activation-aware calibration and introduce output residuals from the dense model as correction terms. We further optimize the solution for efficient GPU execution, enabling scalability to billion-parameter LLMs. Evaluations on LLaMA-2 and LLaMA-3 show that DuoGPT outperforms state-of-the-art structured pruning methods by up to 9.17\% accuracy at an iso-speedup of 1.39$\times$ compared to the baseline dense model. Code is available at GitHub. Ruokai Yin, Yuhang Li 0001, Donghyun Lee 0002, Priyadarshini Panda |
NeurIPS | 1 |
| 2024 | MINT: Multiplier-less INTeger Quantization for Energy Efficient Spiking Neural NetworksabstractWe propose Multiplier-less INTeger (MINT) quantization, a uniform quantization scheme that efficiently compresses weights and membrane potentials in spiking neural networks (SNNs). Unlike previous SNN quantization methods, MINT quantizes memory-intensive membrane potentials to an extremely low precision (2-bit), significantly reducing the memory footprint. MINT also shares the quantization scaling factor between weights and membrane potentials, eliminating the need for multipliers required in conventional uniform quantization. Experimental results show that our method matches the accuracy of full-precision models and other state-of-the-art SNN quantization techniques while surpassing them in memory footprint reduction and hardware cost efficiency at deployment. For example, 2-bit MINT VGG-16 achieves 90.6% accuracy on CIFAR-10, with roughly 93.8% reduction in memory footprint from the full-precision model and 90% reduction in computation energy compared to vanilla uniform quantization at deployment.11Code is available at https://github.com/Intelligent-Computing-Lab-Yale/MINT-Quantization Ruokai Yin, Yuhang Li 0001, Abhishek Moitra, Priyadarshini Panda |
ASPDAC | 1 |
| 2024 | TT-SNN: Tensor Train Decomposition for Efficient Spiking Neural Network TrainingabstractSpiking Neural Networks (SNNs) have gained significant attention as a potentially energy-efficient alternative for standard neural networks with their sparse binary activation. However, SNNs suffer from memory and computation overhead due to spatio-temporal dynamics and multiple backpropagation computations across timesteps during training. To address this issue, we introduce Tensor Train Decomposition for Spiking Neural Networks (TT-SNN), a method that reduces model size through trainable weight decomposition, resulting in reduced storage, FLOPs, and latency. In addition, we propose a parallel computation pipeline as an alternative to the typical sequential tensor computation, which can be flexibly integrated into various existing SNN architectures. To the best of our knowledge, this is the first of its kind application of tensor decomposition in SNNs. We validate our method using both static and dynamic datasets, CIFAR1I0/100 and N-Caltechl0l, respectively. We also propose a TT-SNN-tailored training accelerator to fully harness the parallelism in TT-SNN. Our results demonstrate substantial reductions in parameter size$(7.98\times)$, FLOPs$(9.25\times)$, training time (17.7 %), and training energy (28.3 %) during training for the N-Caltechl0l dataset, with negligible accuracy degradation. Donghyun Lee 0002, Ruokai Yin, Youngeun Kim, Abhishek Moitra, Yuhang Li 0001, Priyadarshini Panda |
DATE | 2 |
| 2024 | Are SNNs Truly Energy-efficient? - A Hardware PerspectiveabstractSpiking Neural Networks (SNNs) have gained attention for their energy-efficient machine learning capabilities, utilizing bio-inspired activation functions and sparse binary spike-data representations. While recent SNN algorithmic advances achieve high accuracy on large-scale computer vision tasks, their energy-efficiency claims rely on certain impractical estimation metrics. This work studies two hardware benchmarking platforms for large-scale SNN inference, namely SATA and SpikeSim. SATA is a sparsity-aware systolic-array accelerator, while SpikeSim evaluates SNNs implemented on In-Memory Computing (IMC) based analog crossbars. Using these tools, we find that the actual energy-efficiency improvements of recent SNN algorithmic works differ significantly from their estimated values due to various hardware bottlenecks. We identify and addresses key roadblocks to efficient SNN deployment on hardware, including repeated computations & data movements over timesteps, neuronal module overhead and vulnerability of SNNs towards crossbar non-idealities. Abhiroop Bhattacharjee, Ruokai Yin, Abhishek Moitra, Priyadarshini Panda |
ICASSP | 2 |
| 2024 | LoAS: Fully Temporal-Parallel Dataflow for Dual-Sparse Spiking Neural NetworksabstractSpiking Neural Networks (SNNs) have gained significant research attention over the past decade due to their potential for enabling resource-constrained edge devices. While existing SNN accelerators efficiently process sparse spikes with dense weights, the opportunities for accelerating SNNs with sparse weights, referred to as dual-sparsity, remain underexplored. In this work, we focus on accelerating dual-sparse SNNs, particularly on their core operation: sparse-matrix-sparse-matrix multiplication (spMspM). Our observations reveal that executing a dual-sparse SNN on existing spMspM accelerators designed for dual-sparse Artificial Neural Networks (ANNs) results in sub-optimal efficiency. The main challenge is that SNNs, which naturally processes multiple timesteps, introducing an additional loop in ANN spMspM, leading to longer latency and more memory traffic. To address this issue, we propose a fully temporal-parallel (FTP) dataflow that minimizes data movement across timesteps and reduces the end-to-end latency of dual-sparse SNNs. To enhance the efficiency of the FTP dataflow, we introduce an FTP-friendly spike compression mechanism that efficiently compresses single-bit spikes and ensures contiguous memory access. Additionally, we propose an FTP-friendly inner-join circuit that reduces the cost of expensive prefix-sum circuits with negligible throughput penalties. These innovations are encapsulated in LoAS, a Low-latency inference Accelerator for dual-sparse SNNs. Running dual-sparse SNN workloads on LoAS demonstrates significant speedup (up to$8.51 \times$) and energy reduction (up to$3.68\times$) compared to prior dual-sparse accelerators. Ruokai Yin, Youngeun Kim, Di Wu 0016, Priyadarshini Panda |
MICRO | 1 |
| 2024 | Do we really need a large number of visual prompts?
Youngeun Kim, Yuhang Li 0001, Abhishek Moitra, Ruokai Yin, Priyadarshini Panda |
Neural Networks | 4 |
| 2023 | Hardware Accelerators for Spiking Neural Networks for Energy-Efficient Edge ComputingabstractNo abstract available. Abhishek Moitra, Ruokai Yin, Priyadarshini Panda |
ACM Great Lakes Symposium on VLSI | 2 |
| 2023 | SATA: Sparsity-Aware Training Accelerator for Spiking Neural NetworksabstractSpiking neural networks (SNNs) have gained huge attention as a potential energy-efficient alternative to conventional artificial neural networks (ANNs) due to their inherent high-sparsity activation. Recently, SNNs with backpropagation through time (BPTT) have achieved a higher accuracy result on image recognition tasks than other SNN training algorithms. Despite the success from the algorithm perspective, prior works neglect the evaluation of the hardware energy overheads of BPTT, due to the lack of a hardware evaluation platform for this SNN training algorithm. Moreover, although SNNs have long been seen as an energy-efficient counterpart of ANNs, a quantitative comparison between the training cost of SNNs and ANNs is missing. To address the aforementioned issues, in this work, we introduce a sparsity-aware training accelerator (SATA), a BPTT-based training accelerator for SNNs. The proposed SATA provides a simple and reconfigurable systolic-based accelerator architecture, which makes it easy to analyze the training energy for BPTT-based SNN training algorithms. By utilizing the sparsity, SATA increases its computation energy efficiency by$5.58\times $compared to the one without using sparsity. Based on SATA, we show quantitative analyses of the energy efficiency of SNN training and make a comparison between the training cost of SNNs and ANNs. The results show that, on Eyeriss-like systolic-based architecture, SNNs consume$1.27\times $more total energy with considering sparsity (spikes, gradient of firing function, and gradient of membrane potential) when compared to ANNs. We find that such high training energy cost is from time-repetitive convolution operations and data movements during backpropagation. Moreover, to propel the future SNN training algorithm design, we provide several observations on energy efficiency for different SNN-specific training parameters and propose an energy estimation framework for SNN training. Ruokai Yin, Abhishek Moitra, Abhiroop Bhattacharjee, Youngeun Kim, Priyadarshini Panda |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | Exploring Lottery Ticket Hypothesis in Spiking Neural Networks
Youngeun Kim, Yuhang Li 0001, Hyoungseob Park, Yeshwanth Venkatesha, Ruokai Yin, Priyadarshini Panda |
ECCV (12) | 5 |
| 2021 | Normalized Stability: A Cross-Level Design Metric for Early Termination in Stochastic ComputingabstractStochastic computing is a statistical computing scheme that represents data as serial bit streams to greatly reduce hardware complexity. The key trade-off is that processing more bits in the streams yields higher computation accuracy at the cost of more latency and energy consumption. To maximize efficiency, it is desirable to account for the error tolerance of applications and terminate stochastic computations early when the result is acceptably accurate. Currently, the stochastic computing community lacks a standard means of measuring a circuit's potential for early termination and predicting at what cycle it would be safe to terminate. To fill this gap, we propose normalized stability, a metric that measures how fast a bit stream converges under a given accuracy budget. Our unit-level experiments show that normalized stability accurately reflects and contrasts the early-termination capabilities of varying stochastic computing units. Furthermore, our application-level experiments on low-density parity-check decoding, machine learning and image processing show that normalized stability can reduce the design space and predict the timing to terminate early. Di Wu 0016, Ruokai Yin, Joshua San Miguel |
ASP-DAC | 2 |
| 2020 | UGEMM: Unary Computing Architecture for GEMM ApplicationsabstractGeneral matrix multiplication (GEMM) is universal in various applications, such as signal processing, machine learning, and computer vision. Conventional GEMM hardware architectures based on binary computing exhibit low area and energy efficiency as they scale due to the spatial nature of number representation and computing. Unary computing, on the other hand, can be performed with extremely simple processing units, often just with a single logic gate. But currently there exist no efficient architectures for unary GEMM. In this paper, we present uGEMM, an area- and energy-efficient unary GEMM architecture enabled by novel arithmetic units. The proposed design relaxes previously-imposed constraints on input bit streams-low correlation and long stream length- and achieves superior area and energy efficiency over existing unary systems. Furthermore, uGEMM's output bit streams exhibit higher accuracy and faster convergence, enabling dynamic energy-accuracy scaling on resource-constrained systems. Di Wu 0016, Ruokai Yin, Hsuan Hsiao, Younghyun Kim 0001, Joshua San Miguel |
ISCA | 3 |