Mohammed Alawad

dblp:131/4987 · DBLP profile ↗
← Back
14ranked-venue papers
9as first author
4since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 8 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Sparsifying Graph Neural Networks with Compressive Sensing
abstract
The computational complexity of graph neural networks (GNNs) presents a significant obstacle to their widespread adoption in various applications. As the size of the input graph increases, the number of parameters in GNN models grows rapidly, leading to increased training and inference times. Existing techniques for reducing GNN complexity, such as train and prune methods and sparse training, often struggle to balance model accuracy with efficiency. In this paper, we address this challenge by proposing a novel approach to sparsify GNNs using sparsity regularization and compressive sensing. By mapping GNN model parameters into a graph and applying sparsity regularization, we induce sparsity in parameter values. Leveraging compressive sensing with Bayesian learning, we identify critical parameters for sparsification, effectively reducing computational costs without sacrificing model accuracy. We evaluate our method on real-world graph datasets and compare it with state-of-the-art techniques. Experimental results demonstrate that our approach achieves higher accuracy while significantly reducing training sparsity and computational requirements, thereby mitigating the impact of large graph sizes on training and inference times. This work sheds light on the potential of compressive sensing for unlocking efficiency in graph-based learning tasks.
Mohammed Alawad, Mohammad Munzurul Islam
ACM Great Lakes Symposium on VLSI1
2024 Probabilistic Bayesian Neural Networks for Efficient Inference
abstract
Bayesian Neural Networks (BNNs) offer a principled framework for modeling uncertainty in deep learning tasks. However, conventional BNNs often suffer from high computational complexity and parameter overhead. In this paper, we propose a novel approach, termed Probabilistic BNN (ProbBNN), which leverages probabilistic computing principles to streamline the inference process. Unlike traditional deterministic approaches, ProbBNN represents inputs and parameters as random variables governed by probability distributions, allowing for the propagation of uncertainty throughout the network. We employ Gaussian Mixture Models (GMMs) to represent the parameters of each neuron or convolutional kernel, enabling efficient encoding and processing of uncertainty. Our approach simplifies the inference process by replacing complex deterministic computations with lightweight probabilistic operations, resulting in reduced computational complexity and improved scalability. Experimental results demonstrate the effectiveness of ProbBNN in achieving competitive accuracy to traditional BNNs while significantly reducing the number of parameters. The transition to ProbBNN yields a reduction of two orders of magnitude in the number of parameters compared to baseline approaches, making our approach promising for deployment in resource-constrained applications such as edge computing and IoT devices.
Mohammed Alawad, Md Ishak
ACM Great Lakes Symposium on VLSI1
2023 Stochastically Pruning Large Language Models Using Sparsity Regularization and Compressive Sensing
abstract
Deep learning models have achieved state-of-the-art performance in many natural language processing tasks. However, their large size and computational complexity make them difficult to deploy in resource-constrained environments. In this paper, we propose a novel approach for reducing the complexity of deep learning models for NLP by combining compressive sensing and Bayesian learning. In our approach, we use compressive sensing with Bayesian learning to identify the most important weights in a deep learning model and represent them in a compressed form. We then prune the non-critical weights and ensure that the model's accuracy is preserved. We evaluate our approach on several NLP tasks, including sentiment analysis and text classification, and compare its performance to that of other compression methods, such as weight pruning and knowledge distillation. Our results show that our approach can significantly reduce the complexity of deep learning models (90% compression) while maintaining high accuracy (<1% drop). In conclusion, our novel approach offers a promising solution for reducing the complexity of deep learning models for NLP and making them more feasible for deployment in resource-constrained environments. By combining compressive sensing and Bayesian learning, we achieve a trade-off between model size and accuracy that is superior to other methods.
Mohammad Munzurul Islam, Mohammed Alawad
ACM Great Lakes Symposium on VLSI2
2023 Node Selection in Federated Learning Using Sparsity Regularization and Compressive Sensing
abstract
In the era of decentralized data and federated learning, node selection plays a pivotal role in ensuring efficient and effective model aggregation while optimizing communication overhead. In this research paper, we propose a novel approach for node selection in federated learning by integrating sparsity regularization and compressive sensing techniques. Federated learning allows training machine learning models on decentralized data sources while preserving privacy and data locality. The node selection process significantly influences model aggregation efficiency and effectiveness. Our approach utilizes sparsity regularization and compressive sensing with Bayesian learning to identify a sparse subset of informative nodes, emphasizing the inclusion of nodes with relevant and non-redundant information for model convergence, thereby reducing communication and computational overhead. Through extensive experimental evaluations on various datasets, we demonstrate significant improvements in model accuracy and convergence rates compared to existing methods, showcasing the remarkable effectiveness of our approach in selecting informative nodes for the success of federated learning applications. The results affirm the potential of our approach in addressing data heterogeneity while advancing efficient, and scalable decentralized machine learning across diverse domains.
Mohammad Munzurul Islam, Mohammed Alawad
ICMLA2
2020 Inter - intra observer variability using deep learning and traditional image processing for breast cancer
abstract
Breast cancer is one of the most common life-threatening diseases that affects women globally. Saudi Arabia is also one of the countries that suffer from a serious number of this disease among women. In terms of diagnosis modalities, a mammogram is the first line for detecting breast cancer. In addition, breast cancer can be screened by real-time ultrasound images, which are of relatively less quality and have more impact (noninvasive) images. Therefore, the purpose of this study is to develop image enhancement techniques using deep learning and image processing techniques. The main goal is to improve the ultrasound images in order to help radiologists screen the disease more accurately. For this study, ninety female patients of ages between 15 - 77 years are considered. These patients were already diagnosed using ultrasound with breast lesions. The images are visually graded and evaluated by two trained radiologists, both pre- and post-enhancement. In particular, two parameters were considered; 1) BI-RAD categories, 2) Breast cancer classification. The agreement between radiologists and post-enhancement was assessed using simple kappa and weighted kappa statistics. Moreover, sensitivity and specificity are also calculated.
Ahmed Almazroa, Barrak Alsomaie, Najd Alluhaydan, Amal Alhaidary, Mohammed Fahim, Wadood Abdul, Mohammed Alawad, Altaf Khan, Ebtihal Alenezi, Taghreed almotairi, AlJowharah Alyahya, Raghad Alfulayj, Manar Althobaiti, Jawaher bin Maythir
DeSE7
2019 Learning Domain Shift in Simulated and Clinical Data: Localizing the Origin of Ventricular Activation From 12-Lead Electrocardiograms
abstract
Building a data-driven model to localize the origin of ventricular activation from 12-lead electrocardiograms (ECG) requires addressing the challenge of large anatomical and physiological variations across individuals. The alternative of a patient-specific model is, however, difficult to implement in clinical practice because the training data must be obtained through invasive procedures. In this paper, we present a novel approach that overcomes this problem of the scarcity of clinical data by transferring the knowledge from a large set of patient-specific simulation data while utilizing domain adaptation to address the discrepancy between the simulation and clinical data. The method that we have developed quantifies non-uniformly distributed simulation errors, which are then incorporated into the process of domain adaptation in the context of both classification and regression. This yields a quantitative model that, with the addition of 12-lead ECG data from each patient, provides progressively improved patient-specific localizations of the origin of ventricular activation. We evaluated the performance of the presented method in localizing 75 pacing sites on three in-vivo premature ventricular contraction (PVC) patients. We found that the presented model showed an improvement in localization accuracy relative to a model trained on clinical ECG data alone or a model trained on combined simulation and clinical data without considering domain shift. Furthermore, we demonstrated the ability of the presented model to improve the real-time prediction of the origin of ventricular activation with each added clinical ECG data, progressively guiding the clinician towards the target site.
Mohammed Alawad
IEEE Trans. Medical Imaging1
2017 Stochastic-Based Multi-stage Streaming Realization of a Deep Convolutional Neural Network (Abstract Only)
Mohammed Alawad, Mingjie Lin
FPGA1
2017 Sketching Computation with Stochastic Processing Engines
abstract
This article explores how to leverage stochastic principles to gracefully exploit partial computation results, hence achieving quality-scalable embedded computing. Our work is inspired by the concept of incremental sketching frequently found in artistic rendering, where the drawing procedure consists of a series of steps, each gradually improving the quality of results. The essence of our approach is to first encode input signals as probability density functions (PDFs), then perform stochastic computing operations on all signals in the probabilistic domain, and finally decode output signals by estimating the PDF of these resulting random samples. Although numerous approximate computing schemes exist, such as inaccurate adders and multipliers that reduce bit width or weaken logic circuit design, none of them can seamlessly improve computing accuracy incrementally without making any changes to the computing hardware at runtime. Furthermore, in conventional embedded computing, a sudden shortage of computing resources, such as premature termination, often means a complete computing failure and totally unusable results. Our sketching computing scheme can readily trade off between the quality of results and computing efforts without modifying its circuit design. To validate our proposed architecture design, we have implemented a proof-of-concept computation sketching engine based on a probabilistic convolver using a Virtex-6 FPGA device. Using three widely deployed image processing applications—image correspondence, image sharpening, and edge detection—we have demonstrated that important embedded computing applications can indeed be “sketched” in a graceful manner using roughly one third the hardware and one fifth the energy compared to the traditional multiplier-based computing method.
Mohammed Alawad, Mingjie Lin
ACM J. Emerg. Technol. Comput. Syst.1
2016 Stochastic-Based Convolutional Networks with Reconfigurable Logic Fabric (Abstract Only)
abstract
Large-scale convolutional neural network (CNN), well-known to be computationally intensive, is a fundamental algorithmic building block in many computer vision and artificial intelligence applications that follow the deep learning principle. This work presents a novel stochastic-based and scalable hardware architecture and circuit design that computes a convolutional neural network with FPGA. The key idea is to implement a multi-dimensional convolution accelerator that leverages the widely-used convolution theorem. Our approach has three advantages. First, it can achieve significantly lower algorithmic complexity for any given accuracy requirement. This computing complexity, when compared with that of conventional multiplierbased and FFT-based architectures, represents a significant performance improvement. Second, this proposed stochastic-based architecture is highly fault-tolerant because the information to be processed is encoded with a large ensemble of random samples. As such, the local perturbations of its computing accuracy will be dissipated globally, thus becoming inconsequential to the final overall results. Overall, being highly scalable and energy efficient, our stochastic-based convolutional neural network architecture is well-suited for a modular vision engine with the goal of performing real-time detection, recognition and segmentation of mega-pixel images, especially those perception-based computing tasks that are inherently fault-tolerant. We also present a performance comparison between FPGA implementations that use deterministic-based and Stochastic-based architectures.
Mohammed Alawad, Mingjie Lin
FPGA1
2015 FIR Filter Based on Stochastic Computing with Reconfigurable Digital Fabric
abstract
FIR filtering is widely used in many important DSP applications in order to achieve filtering stability and linear-phase property. This paper presents a hardware-and energy-efficient approach to implement FIR filtering through reconfigurable stochastic computing. Specifically, we exploit a basic probabilistic principle of summing independent random variables to achieve approximate FIR filtering without costly multiplications. This allows our proposed FIR architecture to achieve about 9 times and 4 times less power consumption than the conventional multiplier-based and DA-based design, respectively. Additionally, when compared with the state-of-the art systolic DA-based design, our design can achieve about 3times reduction in hardware usage.
Mohammed Alawad, Mingjie Lin
FCCM1
2015 Energy-Efficient High-Order FIR Filtering through Reconfigurable Stochastic Processing (Abstract Only)
abstract
High-order FIR filtering is widely used in many important DSP applications in order to achieve filtering stability and linear-phase property. This paper presents a hardware- and energy-efficient approach to implementing energy-efficient high-order FIR filtering through reconfigurable stochastic processing. We exploit a basic probabilistic principle of summing independent random variables to achieve approximate FIR filtering without costly multiplications. Our new multiplierless approach has two distinctive advantages when compared with the conventional multiplier-based or DA-based FIR filtering methods. First, our new probabilistic architecture is especially effective for high-order FIR filtering because it bypasses costly multiplications and does not rely on large size of memory to store store pre-computed coefficient products. Second, this new probabilistic convolver is significantly more robust or fault tolerant than the conventional architecture because all signal values will be represented and computed probabilistically, and local signal corruption can not easily destroy the overall probabilistic patterns, therefore achieving much higher error tolerance. For example, our proposed approach allows our proposed FIR architecture, for a standard 128-tap FIR filter, to achieve about 9 times and 4 times less power consumption than the conventional multiplier-based and DA-based design, respectively. Additionally, when compared with the state-of-the-art systolic DA-based design, our design can achieve about 3 times reduction in hardware usage.
Mohammed Alawad, Mingjie Lin
FPGA1
2014 Energy-efficient multiplier-less discrete convolver through probabilistic domain transformation
abstract
Energy efficiency and algorithmic robustness typically are conflicting circuit characteristics, yet with CMOS technology scaling towards 10-nm feature size, both become critical design metrics simultaneously for modern logic circuits. This paper propose a novel computing scheme hinged on probabilistic domain transformation aiming for both low power operation and fault resilience. In such a computing paradigm, algorithm inputs are first encoded through probabilistic means, which translates the input values into a number of random samples. Subsequently, light-weight operations, such as sim- ple additions will be performed onto these random samples in order to generate new random variables. Finally, the resulting random samples will be decoded probabilistically to give the final results.
Mohammed Alawad, Yu Bai 0004, Ronald F. DeMara, Mingjie Lin
FPGA1
2014 Optimally mitigating BTI-induced FPGA device aging with discriminative voltage scaling (abstract only)
abstract
With the CMOS technology aggressively scaling towards the 22nm node, modern FPGA devices face tremendous aging- induced reliability challenges due to Bias Temperature In- stability (BTI) and Hot Carrier Injection (HCI). This paper presents a novel antiaging technique at logic level that is both scalable and applicable for VLSI digital circuits implemented with FPGA devices. The key idea is to prolong the lifetime of FPGA-mapped designs by strategically elevating the VDD values of some LUTs based on their modular criticality values. Although the idea of scaling VDD in order to improve either energy efficiency or circuit reliability has been explored extensively, our study distinguishes itself by approaching this challenge through analytical procedure, therefore able to maximize the overall reliability of target FPGA design by rigorously modelling the BTI-induce de- vice reliability and optimally solving the VDD assignment problem.
Yu Bai 0004, Mohammed Alawad, Mingjie Lin
FPGA2
2013 Boosting Memory Performance of Many-Core FPGA Device through Dynamic Precedence Graph
abstract
Emerging FPGA device, integrated with abundant RAM blocks and high-performance processor cores, offers an unprecedented opportunity to effectively implement single-chip distributed logic-memory (DLM) architectures [1]. Being “memory-centric”, the DLM architecture can significantly improve the overall performance and energy efficiency of many memory-intensive embedded applications, especially those that exhibit irregular array data access patterns at algorithmic level. However, implementing DLM architecture poses unique challenges to an FPGA designer in terms of 1) organizing and partitioning diverse on-chip memory resources, and 2) orchestrating effective data transmission between on-chip and off-chip memory. In this paper, we offer our solutions to both of these challenges. Specifically, 1) we propose a stochastic memory partitioning scheme based on the well-known simulated annealing algorithm. It obtains memory partitioning solutions that promote parallelized memory accesses by exploring large solution space; 2) we augment the proposed DLM architecture with a reconfigure hardware graph that can dynamically compute precedence relationship between memory partitions, thus effectively exploiting algorithmic level memory parallelism on a per-application basis. We evaluate the effectiveness of our approach (A3) against two other DLM architecture synthesizing methods: an algorithmic-centric reconfigurable computing architectures with a single monolithic memory (A1) and the heterogeneous distributed architectures synthesized according to [1] (A2). To make our comparison fair, in all three architectures, the data path remains the same while local memory architecture differs. For each of ten benchmark applications from SPEC2006 and MiBench [2], we break down the performance benefit of using A3 into two parts: the portion due to stochastic local memory partitioning and the portion due to the dynamic graph-based memory arbitration. All experiments have been conducted with a Virtex-5 (XCV5LX155T-2) FPGA. On average, our experimental results show that our proposed A3 architecture outperforms A2 and A1 by 34% and 250%, respectively. Within the performance improvement of A3 over A2, more than 70% improvement comes from the hardware graph-based memory scheduling.
Yu Bai 0004, Abigail Fuentes-Rivera, Michael Riera, Mohammed Alawad, Mingjie Lin
FCCM4