Shadi Matinizadeh

dblp:374/8383 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Sparse Compressed Quantized Dataflow Architecture (QUDA) for Neuromorphic Computing
abstract
Compressed Sparse Row (CSR) encoding stores only non-zero elements of a sparse matrix along with indexing information for efficient MVM. We propose QUDA, a heterogeneous 2D dataflow architecture that operates directly on CSR-encoded, quantized weights to improve memory efficiency and utilization. It combines flow-through systolic array (FTSA) columns for partial accumulation with a final output-stationary systolic array (OSSA) column for accumulation and spike generation. Our results show significant improvements in resource utilization and power consumption compared to state-of-the-art designs.
Shadi Matinizadeh, Anup Das 0001
FCCM1
2025 Optimizing Memory Latency and Bandwidth of Spiking Neural Network Accelerators on FPGA via Sparse Hashing
abstract
Spiking Neural Networks (SNNs) exploit the natural temporal sparsity by processing discrete events. Introducing additional weight sparsity can further improve their energy efficiency when implemented in hardware. We propose SPARSH, a novel memory organization technique that uses sparse hashing to allow efficient storage and fast access while minimizing memory over-provisioning. Through SPARSH, we make the following four key contributions. First, we introduce an efficient hash function to evenly distribute nonzero weights to buckets that are optimized for an FPGA memory system and placed in a compact memory space to achieve storage efficiency. Second, we use a controller that integrates a Bloom filter to detect sparsity and skip memory accesses for weights and activations that are zero, and for neurons that are in refractory state. This improves memory bandwidth utilization. Third, we propose a scheduler that improves memory latency by prioritizing accesses that hit in the same bucket over other accesses. Finally, we propose an algorithm to select the design parameters of SPARSH based on the sparsity of a target SNN. We implement SPARSH for a recent SNN accelerator on a Virtex UltraScale and evaluate using seven SNNs models. We show that SPARSH reduces memory over-provisioning by 3.2×, access latency by 77%, and bandwidth utilization by 68% with a marginal increase in resource utilization.
Shadi Matinizadeh, M. Lakshmi Varshika, Anup Das 0001
ICCAD1
2025 A Digital Neuromorphic Architecture for Unsupervised Shortest Path Computation on Real-World Graphs
abstract
Graphs are popular tools for analyzing interconnected data entities. We propose SENTIENCE, a novel approach to computing the shortest path in a graph in an unsupervised manner drawing inspiration from the hippocampus. SENTIENCE uses laterally-connected neurons to represent nodes and synapses to represent edges. The strength of a synaptic connection is encoded as the axonal delay. When a neuron (source) is excited, a wavefront of neural activity is created that propagates through the graph via the connected nodes. SENTIENCE uses the Eligibility Propagation (E-Prop) algorithm to learn the sequence of wave movement through the shortest path in a graph, which can be obtained by identifying earliest firing (eligible) neighbors by tracing back in time from when the wavefront reaches a destination. We introduce a lightweight hardware design for the axonal plasticity and the E-Prop learning mechanism of SENTIENCE. We propose a tile-based architecture to address the scalability of SENTIENCE for real-world graphs with irregular data dependency. We evaluate SENTIENCE on a Versal VPK 180 FPGA and show that SENTIENCE consumes on average 10× less resources and 12× less power compared to a state-of-the-art.
Arghavan Mohammadhassani, Shadi Matinizadeh, M. Lakshmi Varshika, Anup Das 0001
ISCAS2
2024 An Open-Source And Extensible Framework for Fast Prototyping and Benchmarking of Spiking Neural Network Hardware
abstract
Spiking neural networks (SNNs) are bioplausible machine learning models that use discrete spikes to encode, compute, and transmit information. Combined with event-driven low-power hardware, SNNs can improve the energy efficiency of learning tasks. Although there have been several efforts to build SNN hardware, there is no uniform framework to verify and benchmark these designs in terms of key hardware performance metrics such as inference accuracy, area, power consumption, and throughput. We propose PRONTO, an open-source and extensible framework to verify SNN hardware for different learning tasks and datasets. Given the ubiquity of PyTorch in the machine learning community and for demonstration purposes, the frontend of PRONTO is integrated with a torch-based SNN simulator for model specification and training. Its backend is integrated with an open-source quantized SNN hardware. PRONTO interfaces with a torch code to generate input stimuli which are then driven to SNN hardware through a configurable SystemVerilog testbench, verifying the design across various SNN-specific configurations. PRONTO utilizes a dataflow-based approach to validate SNN models that are segmented and run on a mix of software and hardware platforms. We describe PRONTO and evaluate it using six datasets spanning image, audio, and text classification. We present benchmark results for various input settings. PRONTO is available under an open-source licensing to provide a platform to evaluate all current and future SNN hardware designs. We believe PRONTO will substantially reduce the design verification effort, thus facilitating fast design prototyping.
Shadi Matinizadeh, Anup Das 0001
FPL1