Faaiz Asim

dblp:309/4391 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 FlowQ: Fixed-point Low-precision Post-Training Quantization Framework for Efficient and Accurate SNN Inference
abstract
We propose FlowQ, a post-training quantization (PTQ) framework for spiking neural networks (SNNs) that balances accuracy and hardware-efficiency through quantizer design and calibration-based scale optimization. While using different scales for weights and membrane potentials preserves accuracy, it typically incurs high hardware cost. In contrast, shared scale factors reduce hardware complexity but lead to significant accuracy degradation. FlowQ bridges this gap by using a hardware-friendly quantizer with different scales that differ by a power-of-two, allowing multiplications to be replaced with simple bit-shift operations for negligible overhead. To further improve accuracy, we present FlowTune, a calibration algorithm that iteratively optimizes FlowQ’s scale factors by minimizing mean-squared error, outperforming the commonly used absolute max-based scaling in SNN PTQ. Extensive experiments on CIFAR-10, CIFAR-100, DVS-Gestures, and ImageNet demonstrate the effectiveness of our approach. For example, on CIFAR-10 with VGG-16 at 4-bit precision, FlowQ achieves only a 1.48% accuracy drop. Compared to shared scale factors, FlowQ improves accuracy by 77.4% with just 1% energy and 0.25% area overhead.
Faaiz Asim, Sanhtet Aung, Jongeun Lee
ASP-DAC1
2026 AccelOrb: FPGA Acceleration of Orb v2 for Fast Molecular Dynamics
abstract
Machine learning force fields (MLFFs) offer high accuracy at reduced computational cost, but their repeated evaluation at every molecular dynamics (MD) timestep leads to prohibitive runtime, limiting simulations to short physical timescales. In this work, we accelerate Orb v2, a Graph Neural Network–based MLFF, by targeting its dominant Attention Interaction Network (AIN), which accounts for approximately 97% of inference time. We identify on-chip memory constraints and a latency-memory trade-off in block-based computation as the key challenges to efficient acceleration. To address these challenges, we propose AccelOrb, a memory-aware FPGA acceleration of AIN based on a single streaming kernel that employs a non-uniform tiling strategy to balance latency and memory footprint, and selectively maps parameters beyond BRAM. Implemented on a resource-constrained AMD Alveo U50 FPGA, our approach achieves 9.19× and 8.18× speedup over CPU and GPU baselines while operating within tight power and resource budgets.
Sunjae Kim, Gwanhong Park, Jeawoo Lim, Faaiz Asim, Jongeun Lee
FCCM4
2024 Extending Neural Processing Unit and Compiler for Advanced Binarized Neural Networks
abstract
Binarized neural networks (BNNs) are one of the most promising approaches to deploy deep neural network models on resource-constrained devices. However, there is very little support on compilers and programmable accelerators for BNNs especially with the modern BNNs that use scale factors and skip connections to maximize network performance. In this paper we present a set of methods to extend a neural processing unit (NPU) and a compiler to support modern BNNs. Our novel ideas include (i) batch-norm folding for binarized layers with scale factors and skip connections, (ii) efficient handling of convolutions with few input channels, and (iii) bit-packing pipelining. Our evaluation using BiRealNet-18 on an FPGA board demonstrates that our compiler-architecture hybrid approach can yield significant speedups for binary convolution layers over the baseline NPU. Also our approach gives 3.6~5.5 $\times$ better end-to-end performance on BiRealNet-18 compared with previous BNN compiler approaches.
Minjoon Song, Faaiz Asim, Jongeun Lee
ASPDAC2
2023 Partial Sum Quantization for Reducing ADC Size in ReRAM-Based Neural Network Accelerators
abstract
While resistive random-access memory (ReRAM) crossbar arrays have the potential to significantly accelerate deep neural network (DNN) training through fast and low-cost matrix–vector multiplication, peripheral circuits like analog-to-digital converters (ADCs) create a high overhead. These ADCs consume over half of the chip power and a considerable portion of the chip cost. To address this challenge, we propose advanced quantization techniques that can significantly reduce the ADC overhead of ReRAM crossbar arrays (RCAs). Our methodology interprets ADC as a quantization mechanism, allowing us to scale the range of ADC input optimally along with the weight parameters of a DNN, resulting in multiple-bit reduction in ADC precision. This approach reduces ADC size and power consumption by several times, and it is applicable to any DNN type (binarized or multibit) and any RCA size. Additionally, we propose ways to minimize the overhead of the digital scaler, which is an essential part of our scheme and sometimes required. Our experimental results using ResNet-18 on the ImageNet dataset demonstrate that our method can reduce the size of the ADC by 32 times compared to ISAAC with only a minimal accuracy loss degradation of 0.24%. We also present evaluation results in the presence of ReRAM nonideality (such as stuck-at fault).
Azat Azamat, Faaiz Asim, Jongeun Lee
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2022 Centered Symmetric Quantization for Hardware-Efficient Low-Bit Neural Networks
Faaiz Asim, Jaewoo Park 0006, Azat Azamat, Jongeun Lee
BMVC1
2021 Quarry: Quantization-based ADC Reduction for ReRAM-based Deep Neural Network Accelerators
abstract
ReRAM (Resistive Random-Access Memory) crossbar arrays have the potential to provide extremely fast and low-cost DNN (Deep Neural Network) acceleration. However, peripheral circuits, in particular ADCs (Analog-Digital Converters), can be a large overhead and/or slow down the operation considerably. In this paper we propose to use advanced quantization techniques to reduce the ADC overhead of ReRAM crossbar arrays. Our method does not require any hardware change but can reduce the overhead of ADC greatly. Our methodology is also general, having no restriction in terms of DNN type (binarized or multi-bit) or ReRAM crossbar array size. Our experimental results using ResNet on ImageNet dataset demonstrate that our method can reduce the size of ADC by 32× compared with ISAAC at very little accuracy loss of 0.24%p.
Azat Azamat, Faaiz Asim, Jongeun Lee
ICCAD2