Sajjad Moazeni

dblp:145/5401 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
6since 2021 · last 2024
0000-0003-1819-0714ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 5 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Pseudo-Sim: An Accurate Analytical Modeling Framework for Systolic Array Architectures
abstract
Array-based accelerators have emerged as powerful compute engines for meeting the immense computational de-mands of artificial intelligence (AI). However, current tools for optimizing their complex multidimensional hardware-software design space feature a strict tradeoff between accuracy and runtime. Here, we introduce Pseudo-Sim, a rapid and accurate methodology to simulate key system parameters including the number of compute clock cycles, off-chip DRAM accesses, and SRAM accesses. We validate the effectiveness of our tool by running it on the Resnet50V1.5 and BERT-Base neural networks for multiple hardware configurations and comparing results to those from a a cycle-accurate simulator. Pseudo-Sim achieves$10^{4}\times$speedup).
Daniel Sturm 0003, Sajjad Moazeni
ICCD2
2024 Accelerating Cascade Classifier Training with Genetic Algorithms for Edge ML Applications
abstract
Object detection is a crucial task in computer vision with applications spanning from face recognition to autonomous driving. While today’s CNN-based methods have shown great success for this aim, their relatively large model size and computational complexity limit their use for edge devices such as micro-controllers. On the other hand, the Viola-Jones algorithm has long been a cornerstone in this field, offering robustness and accuracy. This algorithm can be more compact than even compressed CNN models. However, as datasets and feature spaces grow, the computational demands of training an Adaboost classifier can become prohibitively high. This can be a bottleneck where training of the model needs to also be performed at the edge for cost and privacy reasons. In this paper, we present an innovative approach to address this challenge by incorporating Genetic Algorithms (GA) and LightGBM into the Adaboost framework for efficient feature selection, reducing training time by a factor of $50 \times$ without sacrificing accuracy. Additionally, our model exhibits a significantly lower memory footprint, with a size of 20 kB compared to a compressed CNN-based YOLOX architecture with a model size of 314 kB. This makes our approach particularly suitable for detection in edge devices and TinyML community.
Abhishek Saini, Sajjad Moazeni
ICIP2
2024 A Mixed-Signal Compute-in-Memory Architecture for Solving All-to-All Connected MAXCUT Problems with Sub-µs Time-to-Solution
abstract
Combinatorial and discrete optimization problems are prevalent in fields such as artificial intelligence, supply chain management, and wireless communications. The Ising machine, a quantum-inspired paradigm, offers a novel approach to accelerate these computations. However, realizing an Ising machine in an area/energy-efficient and scalable manner with low compute latency in CMOS is challenging. In this work, we propose a new mixed-signal SRAM-based compute-in-memory architecture to perform as a simulated bifurcation (SB) Ising machine. This realization leverages the inherent noise of analog computing, accelerating the time to anneal by injecting decaying noise in the analog domain. We have verified our solution and studied parameter optimization in the 180nm CMOS process with up to 60 spins using a post-layout co-simulation framework based on Synopsys PrimeSim. We benchmarked our design on 60-node random binary MAXCUT problems with all-to-all connections. This architecture achieves +95% of the ground state consistently over 10 graphs with < 1µs run time and 7.6mW average power.
Alana Dee, Katherine Bennett, Sajjad Moazeni
ISCAS3
2024 OFHE: An Electro-Optical Accelerator for Discretized TFHE
abstract
This paper presents OFHE, an electro-optical accelerator designed to process Discretized TFHE (DTFHE) operations, which encrypt multi-bit messages and support homomorphic multiplications, lookup table operations and full-domain functional bootstrappings. While DTFHE is more efficient and versatile than other fully homomorphic encryption schemes, it requires 32-, 64-, and 128-bit polynomial multiplications, which can be time-consuming. Existing TFHE accelerators are not easily upgradable to support DTFHE operations due to limited datapaths, a lack of datapath bit-width reconfigurability, and power inefficiencies when processing FFT and inverse FFT (IFFT) kernels. Compared to prior TFHE accelerators, OFHE addresses these challenges by improving the DTFHE operation latency by 8.7%, the DTFHE operation throughput by 57%, and the DTFHE operation throughput per Watt by 94%.
Mengxin Zheng, Cheng Chu, Qian Lou, Nathan Youngblood, Sajjad Moazeni, Lei Jiang 0001
ISLPED6
2024 Design of a Mixed-Signal Compute-in-Memory Ising Solver With Sub-μs Time-to-Solution and Optimal Decaying Noise Profile
abstract
Combinatorial and discrete optimization problems are prevalent in fields such as artificial intelligence, supply chain management, and wireless communications. The Ising machine, a quantum-inspired paradigm, offers a novel approach to accelerate these computations. However, realizing an Ising machine in an area/energy-efficient and scalable manner with low compute latency in CMOS is challenging. In this work, we propose a new mixed-signal SRAM-based compute-in-memory architecture to perform as a simulated bifurcation (SB) Ising machine. This realization leverages the inherent noise of analog computing, accelerating the time to anneal by injecting decaying noise in the analog domain. We have verified our solution and studied parameter optimization in the 180nm CMOS process with up to 60 spins using a post-layout co-simulation framework based on Synopsys PrimeSim. We benchmarked our design on 60-node random binary MAXCUT problems with all-to-all connections. This architecture achieves +95% of the ground state consistently over 10 graphs with$\lt {1\mu s}$run time and 7.6mW average power. This work further discusses techniques to optimize the injected decaying noise profile and SB tuning parameters, which are both crucial to high accuracy bifurcation. We analyze the performance of our proposed tuning methods over different graph densities, increasing number of nodes, and multi-bit graph edge weights.
Alana Dee, Duy Vuong, Katherine Bennett, Sajjad Moazeni
IEEE Trans. Circuits Syst. I Regul. Pap.4
2023 Scalable Coherent Optical Crossbar Architecture using PCM for AI Acceleration
abstract
Optical computing has recently been proposed as a new compute paradigm to meet the demands of future AI/ML workloads in datacenters and supercomputers. However, proposed implementations so far suffer from lack of scalability, large footprints and high power consumption, and incomplete system-level architectures inhibit integration within existing datacenter systems for real-world applications. In this work, we present a truly scalable optical AI accelerator based on a crossbar architecture. We have considered all major roadblocks and address them in this design. Weights will be stored on-chip using phase change material (PCM) that can be monolithically integrated in silicon photonic processes. All electro-optical components and circuit blocks are modeled based on measured performance metrics in a 45nm monolithic silicon photonic process, which can be co-packaged with advanced CPU/GPUs and HBM memories. We also present a system-level modeling and analysis of our chip's performance for the Resnet-50V1.5, considering all critical parameters, including memory size, array size, photonic losses, and energy consumption of peripheral electronics. Both on-chip SRAM and off-chip DRAM energy overheads have been considered in this modeling. We additionally address how using a dual-core crossbar design can eliminate programming time overhead at practical SRAM block sizes and batch sizes. Our results show that a 128 × 128 proposed architecture can achieve inference per second (IPS) similar to Nvidia A100 GPU at 15.4× lower power and 7.24× lower area.
Daniel Sturm 0003, Sajjad Moazeni
DATE2
2014 Generalisation of code division multiple access systems and derivation of new bounds for the sum capacity
abstract
In this study, the authors explore a generalised scheme for the synchronous code division multiple access (CDMA). In this scheme, unlike the standard CDMA systems, each user has different codewords for communicating different messages. Two main problems are investigated. The first problem concerns whether uniquely detectable overloaded matrices (an injective matrix, i.e. the inputs and outputs are in one‐to‐one correspondence depending on the input alphabets) exist in the absence of additive noise, and if so, whether there are any practical optimum detectors for such input codewords. The second problem is about finding tight bounds for the sum channel capacity. In response to the first problem, the authors have constructed uniquely detectable matrices for the generalised scheme and the authors have developed practical maximum likelihood detection algorithms for such codes. In response to the second problem, lower bounds and conjectured upper bounds are derived. The results of this study are superior to other standard overloaded CDMA codes since the generalisation can support more users than the previous schemes.
Shayan Dashmiz, Mohammad Reza Takapoui, Sajjad Moazeni, Mehrdad Moharrami, Melika Abolhasani, Farrokh Marvasti
IET Commun.3