Hassaan Saadat

dblp:207/7217 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
5since 2021 · last 2025
0000-0003-3691-4130ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 5 first-author · 5 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 A Novel Covert Timing Channel for Cloud FPGAs
abstract
This paper presents a novel covert timing channel (CTC) that enables a malicious entity to exfiltrate data from a benign cloud FPGA user without requiring dedicated outgoing messages from the cloud FPGA, minimizing the detection risk by both the victim and the cloud service provider. The proposed CTC exploits the handshake signals of the Advanced eXtensible Interface (AXI) protocol and interpacket delay of the Internet to establish the CTC from a cloud FieldProgrammable Gate Array (FPGA) to an off-cloud computer. This paper analyzes the bit-error rate (BER) of the AXI-based CTC under varying conditions and demonstrates its effectiveness in truly enabling remote power analysis attacks on cloud services, such as Amazon Web Services Elastic Compute Cloud (AWS EC2). The proposed CTC achieves a BER as low as 0.01988%.
Brian Udugama, Darshana Jayasinghe, Hassaan Saadat, Aleksandar Ignjatovic, Sri Parameswaran
DAC3
2024 SEA: Sign-Separated Accumulation Scheme for Resource-Efficient DNN Accelerators
abstract
Deep neural network (DNN) accelerators targeting training need to support the resource-hungry floating-point (FP) arithmetic. Typically, the additions (accumulations) are performed with higher precision compared to multiplications, resulting in a substantial proportion of resources consumed by the adders. In this paper, we aim to improve the resource efficiency of FP addition in DNN training accelerators and present a sign-separated accumulation scheme (SEA). The proposed SEA scheme accumulates the same-signed terms separately, followed by a final addition of the two oppositely-signed sub-accumulations. The separate sub-accumulations, when performed using our resource-efficient same-signed FP adders, result in significant improvement in overall resource efficiency. We present a SEA-based systolic array design, involving a novel processing element design, capable of performing FP matrix multiplications for DNN training. Our experimental results show substantial improvements in delay, area-delay product (ADP), and energy consumption, by achieving reductions of up to 19.8% in delay, 29.1% in ADP and 30.7% in energy, across various adder-multiplier combinations compared to the original designs. SEA introduces no approximations in accumulation or modifications in the rounding mechanism. We integrate the SEA-based systolic array into the open-source Gemmini [1] ecosystem for use by the broader community.
Hassaan Saadat, Haris Javaid, Hasindu Gamaarachchi, David S. Taubman, Sri Parameswaran
DATE2
2024 Accelerating Chaining in Genomic Analysis Using RISC- V Custom Instructions
abstract
This paper presents a method for designing custom instructions tailored to RISC-V processors, focusing on optimizing the chaining step of Minimap2 (a software tool used to analyze DNA data emanating from third-generation sequencing machines). This custom instruction design involves employing an architectural template within the Rocket Custom Coprocessor (RoCC) unit of Rocket Chip, an open-source hardware implementation of RISC- VISA, aided by a heuristic algorithm that facilitates extracting custom instructions from high-level C code targeting the proposed architectural template. Two types of instructions are created in this work: complex computational instructions; and instructions that load static data apriori so that these data are not repeatedly brought in from the memory. The resulting custom instructions integrated into Rocket Chip demonstrate a speedup of up to 2.4 × in the chaining step of Minimap2 with no adverse impact on the final mapping accuracy compared to the original software. The acceleration of Minimap2 's chaining stage on a RISC-V processor enhances its portability and energy efficiency, making third-generation DNA sequence analysis more accessible in various settings.
Kisaru Liyanage, Hasindu Gamaarachchi, Hassaan Saadat, Tuo Li 0001, Hiruna Samarakoon, Sri Parameswaran
DATE3
2023 ApproxTrain: Fast Simulation of Approximate Multipliers for DNN Training and Inference
abstract
Edge training of deep neural networks (DNNs) is a desirable goal for continuous learning; however, it is hindered by the enormous computational power required by training. Hardware approximate multipliers have shown their effectiveness in gaining resource efficiency in DNN inference accelerators; however, training with approximate multipliers is largely unexplored. To build resource-efficient accelerators with approximate multipliers supporting DNN training, a thorough evaluation of training convergence and accuracy for different DNN architectures and different approximate multipliers is needed. This article presents ApproxTrain, an open-source framework that allows fast evaluation of DNN training and inference using simulated approximate multipliers. ApproxTrain is as user-friendly as TensorFlow (TF) and requires only a high-level description of a DNN architecture along with C/C++ functional models of the approximate multiplier. We improve the speed of the simulation at the multiplier level by using a novel LUT-based approximate floating-point (FP) multiplier simulator on GPU (AMSim). Additionally, a novel flow is presented to seamlessly convert C/C++ functional models of approximate FP multipliers into AMSim. ApproxTrain leverages CUDA and efficiently integrates AMSim into the TensorFlow library to overcome the absence of native hardware approximate multiplier in commercial GPUs. We use ApproxTrain to evaluate the convergence and accuracy performance of DNN training with approximate multipliers for three application domains: image classification, object detection, and neural machine translation. The evaluations demonstrate similar convergence behavior and negligible change in test accuracy compared to FP32 and Bfloat16 multipliers. Compared to CPU-based approximate multiplier simulations in training and inference, the GPU-accelerated ApproxTrain is more than$2500\times $faster. Based on highly optimized closed-source cuDNN/cuBLAS libraries with native hardware multipliers, the original TensorFlow is, on average, only$8\times $faster than ApproxTrain.
Hassaan Saadat, Hasindu Gamaarachchi, Haris Javaid, Xiaobo Sharon Hu, Sri Parameswaran
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 Approximate Computing for ML: State-of-the-art, Challenges and Visions
abstract
In this paper, we present our state-of-the-art approximate techniques that cover the main pillars of approximate computing research. Our analysis considers both static and reconfigurable approximation techniques as well as operation-specific approximate components (e.g., multipliers) and generalized approximate highlevel synthesis approaches. As our application target, we discuss the improvements that such techniques bring on machine learning and neural networks. In addition to the conventionally analyzed performance and energy gains, we also evaluate the improvements that approximate computing brings in the operating temperature.
Georgios Zervakis 0001, Hassaan Saadat, Hussam Amrouch, Andreas Gerstlauer, Sri Parameswaran, Jörg Henkel
ASP-DAC2
2020 WEID: Worst-case Error Improvement in Approximate Dividers
abstract
Approximate integer dividers suffer from unreasonably high worst-case relative errors (such as 50% or 100%), which can adversely affect the application-level output. In this paper, we propose WEID, which is a novel lightweight method to improve the worst-case relative errors in approximate integer dividers. We first present an in-depth analysis to gain insights into the cause of the high worst-case relative error. Based on our insights, we propose a novel method to detect when an error occurs in an approximate divider, and modify the output to reduce the error. Further, we present the hardware realization of WEID method and demonstrate that it can be generically coupled with several state-of-the-art approximate dividers. Our results show that for 32-by-16 dividers, WEID reduces worstcase relative errors from 100% to ~20%, while still achieving ~80% and ~70% reduction in delay and energy compared to an accurate array divider.
Hassaan Saadat, Haris Javaid, Aleksandar Ignjatovic, Sri Parameswaran
ASP-DAC1
2020 REALM: Reduced-Error Approximate Log-based Integer Multiplier
abstract
We propose a new error-configurable approximate unsigned integer multiplier named REALM. It incorporates a novel error-reduction method into the classical approximate log-based multiplier. Each power-of-two-interval of the input operands is partitioned into M×M segments, and an error-reduction factor for each segment is analytically determined. These error-reduction factors can be used across any power-of-two-interval, so we quantize only M2factors and store them in the form of read-only hardwired lookup tables to keep the resource overhead to a minimum. Error characterization of REALM shows that it achieves very low error bias (mostly ≤0.05%), along with lower mean error (from 0.4% to 1.6%), and lower peak error (from 2.08% to 7.4%) than the classical approximate log-based multiplier and its state-of-the-art derivatives (mean errors ≥2.6% and peak errors ≥7.8%). Synthesis results using TSMC 45nm standard-cell library show that REALM enables significant power-efficiency (66% to 86% reduction) and area-efficiency (50% to 76% reduction) when compared with the accurate integer multiplier. We show that REALM produces Pareto optimal design trade-offs in the design space of state-of-the-art approximate multipliers. Application-level evaluation of REALM demonstrates that it has negligible effect on the output quality.
Hassaan Saadat, Haris Javaid, Aleksandar Ignjatovic, Sri Parameswaran
DATE1
2019 Approximate Integer and Floating-Point Dividers with Near-Zero Error Bias
abstract
We propose approximate dividers with near-zero error bias for both integer and floating-point numbers. The integer divider, INZeD, is designed using a novel, analytically deduced error-correction method in an approximate log based divider. The floating-point divider, FaNZeD, is based on a highly optimized mantissa divider that is inspired by INZeD. Both of the dividers are error configurable.
Hassaan Saadat, Haris Javaid, Sri Parameswaran
DAC1
2018 Minimally Biased Multipliers for Approximate Integer and Floating-Point Multiplication
abstract
Approximate multipliers enable the saving of area and power for implementation of many modern error-resilient compute-intensive applications. In this paper, we first propose a novel error-configurable minimally biased approximate integer multiplier (MBM) design. The proposed MBM design is devised by coupling a unique error-reduction mechanism with an approximate log based integer multiplier. Next, we propose an optimization (by removing leading-one detection and barrel shifting logic) of the MBM and a class of state-of-the-art approximate integer multipliers (DRUM and SSM), so that they can be efficiently used in approximate floating-point (FP) multipliers. Then, we propose a set of new approximate FP multipliers and we show that these FP multipliers lie on the Pareto front on the design spaces of area versus error and power versus error. We synthesize the designs using the TSMC 45-nm standard-cell library. Results show that the MBM integer design offers optimal points in the design space, offering up to 75% area reduction and 84% power reduction with <;0.1% error bias, when compared with the accurate version. The proposed approximate FP multipliers offer better error-efficiency tradeoffs than traditional precision scaling. The FP design space can offer up to 57x power and 28x area improvement for less than 25% peak error, 7% mean error, and 4% error bias, when compared with the IEEE-754 single-precision FP multiplier. We also perform application-level evaluations of the proposed approximate integer and FP multipliers, showing that our proposed multipliers enable significant power and area reduction with minimal degradation in applications' output quality.
Hassaan Saadat, Haseeb Bokhari, Sri Parameswaran
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2017 Hardware approximate computing: how, why, when and where? (special session)
abstract
Approximate computing in hardware is generally aimed at power or energy optimization as the primary target. We suggest that hardware approximate computing can be more beneficial when area reduction is the primary target. Additionally, we advocate that the hardware approximation schemes which allow usage of high-level libraries for their sub-components can leverage the power offered by modern synthesis tools. We demonstrate using experimental results that such approximations therefore achieve more efficient synthesis than the deeply hierarchical approximations.
Hassaan Saadat, Sri Parameswaran
CASES1