Md. Al Maruf

dblp:233/2115 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
5since 2021 · last 2024
0000-0001-5752-8939ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Optimizing DNN training with pipeline model parallelism for enhanced performance in embedded systems
abstract
Deep Neural Networks (DNNs) have gained widespread popularity in different domain applications due to their dominant performance. Despite the prevalence of massively parallel multi-core processor architectures, adopting large DNN models in embedded systems remains challenging, as most embedded applications are designed with single-core processors in mind. This limits DNN adoption in embedded systems due to inefficient leveraging of model parallelization and workload partitioning. Prior solutions attempt to address these challenges using data and model parallelism. However, they lack in finding optimal DNN model partitions and distributing them efficiently to achieve improved performance. This paper proposes a DNN model parallelism framework to accelerate model training by finding the optimal number of model partitions and resource provisions. The proposed framework combines data and model parallelism techniques to optimize the parallel processing of DNNs for embedded applications. In addition, it implements the pipeline execution of the partitioned models and integrates a task controller to manage the computing resources. The experimental results for image object detection demonstrate the applicability of our proposed framework in estimating the latest execution time and reducing overall model training time by almost 44.87% compared to the baseline AlexNet convolutional neural network (CNN) model.
Md. Al Maruf, Akramul Azim, Nitin Auluck, Mansi Sahi
J. Parallel Distributed Comput.1
2024 Dynamic hierarchical intrusion detection task offloading in IoT edge networks
abstract
Abstract The Internet of Things (IoT) has gained widespread importance in recent time. However, the related issues of security and privacy persist in such IoT networks. Owing to device limitations in terms of computational power and storage, standard protection approaches cannot be deployed. In this article, we propose a lightweight distributed intrusion detection system (IDS) framework, called FCAFE‐BNET (Fog based Context Aware Feature Extraction using BranchyNET). The proposed FCAFE‐BNET approach considers versatile network conditions, such as varying bandwidths and data loads, while allocating inference tasks to cloud/edge resources. FCAFE‐BNET is able to adjust to dynamic network conditions. This can be advantageous for applications with particular quality of service requirements, such as video streaming or real‐time communication, ensuring a steady and reliable performance. Early exit deep neural networks (DNNs) have been employed for faster inference generation at the edge. Often, the weights that the model learns in the initial layer may be sufficiently qualified to perform the required classification tasks. Instead of using subsequent layers of DNNs for generating the inference, we have employed the early‐exit mechanism in the DNNs. Such DNNs help to predict a wide range of testing samples through these early‐exit branches, upon crossing a threshold. This method maintains the confidence values corresponding to the inference. Employing this approach, we achieved a faster inference, with significantly high accuracy. Comparative studies exploit manual feature extraction techniques, that can potentially overlook certain valuable patterns, thus degrading classification performance. The proposed framework converts textual/tabular data into 2‐D images, allowing the DNN model to autonomously learns its own features. This conversion scheme facilitated the identification of various intrusion types, ranging from 5 to 14 different categories. FCAFE‐BNET works for both network‐based and host‐based IDS: NIDS and HIDS. Our experiments demonstrate that, in comparison with recent approaches, FCAFE‐BNET achieves a 39.12%–50.23% reduction in the total inference time on benchmark real‐world datasets, such as: NSL‐KDD, UNSW‐NB 15, ToN_IoT, and ADFA_LD.
Mansi Sahi, Nitin Auluck, Akramul Azim, Md. Al Maruf
Softw. Pract. Exp.4
2023 Towards Safe Online Machine Learning Model Training and Inference on Edge Networks
abstract
With the increasing demand for edge computing in cyber-physical system (CPS) applications, ensuring the safety and reliability of machine learning models running on edge devices during online model training and inference is essential. Although data and model parallelism offer significant advantages for large machine learning model training, adopting parallel computing architecture in edge networks is challenging. It introduces safety concerns while splitting and integrating machine learning models over different computing nodes, which can pose risks to the integrity and reliability of the system. Therefore, online model training and inference in edge networks require a safe parallel computing architecture to achieve improved performance with optimal resource utilization. To address this challenge, we propose an efficient machine learning model partitioning algorithm that considers the safety constraint and requirements of edge networks and includes the triple-modular redundancy (TMR) technique for trusted computation. Our proposed approach achieves a significant speedup of approximately 56.3% in net training time compared to the non-partitioning approach, making it more efficient and suitable for real-time applications in edge networks.
Md. Al Maruf, Akramul Azim, Nitin Auluck, Mansi Sahi
ICMLA1
2022 Faster Fog Computing Based Over-the-Air Vehicular Updates: A Transfer Learning Approach
abstract
Fog computing is a promising option for time sensitive vehicular over-the-air (OTA) updates, as it can offer enhanced network durability and lower communication delays, as compared to the cloud. Fog node utilization for updates is non-deterministic, largely owing to the patterns in vehicular traffic. The resultant over provisioning of resources manifests itself in increased communication and handover delays. Based on an analysis of the regional traffic pattern for a particular time period, our proposed algorithm determines the optimal number of fog nodes required for OTA updates. In order to pinpoint the traffic load and perform fog node distribution, we employ k-means clustering. The efficacy of our proposed approach is demonstrated using a case study that considers handover delay, propagation delay, transmission rate and vehicular mobility to predict the OTA update time. We employ a machine learning model for predicting the communication delay between fog devices and vehicles. Using the European WiFi hotspot signal strength NYC dataset and the 5G dataset, we observe that the proposed approach increases the net reserve fog resources by 26.57 percent on an average, and reduces the OTA update time by 5.34 percent. We test the scalability of the proposed approach by analyzing the performance in terms of average throughput while varying the number of vehicles and OTA update size. We observe that a system with less traffic and small update size overall delivers a higher average throughput of 46 Mbps versus one with more traffic and large update size overall, which provides an average throughput of 30 Mbps. The performance of the proposed OTA update scheme on simulations has been corroborated by implementation on a real-world testbed.
Md. Al Maruf, Anil Singh, Akramul Azim, Nitin Auluck
IEEE Trans. Serv. Comput.1
2021 A Framework for Partitioning Support Vector Machine Models on Edge Architectures
abstract
Current IoT applications generate huge volumes of complex data that requires agile analysis in order to obtain deep insights, often by applying Machine Learning (ML) techniques. Support vector machine (SVM) is one such ML technique that has been used in object detection, image classification, text categorization and Pattern Recognition. However, training even a simple SVM model on big data takes a significant amount of computational time. Due to this, the model is unable to react and adapt in real-time. There is an urgent need to speedup the training process. Since organizations typically use the cloud for this data processing, accelerating the training process has the advantage of bringing down costs. In this paper, we propose a model partitioning approach that partitions the tasks of Stochastic Gradient Descent based Support Vector Machines (SGD-SVM) on various edge devices for concurrent computation, thus reducing the training time significantly. The proposed partitioning mechanism not only brings down the training time but also maintains the approximate accuracy over the centralized cloud approach. With a goal of developing a smart objection detection system, we conduct experiments to evaluate the performance of the proposed method using SGD-SVM on an edge based architecture. The results illustrate that the proposed approach significantly reduces the training time by 47%, while decreasing the accuracy by 2%, and offering an optimal number of partitions.
Mansi Sahi, Md. Al Maruf, Akramul Azim, Nitin Auluck
SMARTCOMP2
2020 Mushroom Demand Prediction Using Machine Learning Algorithms
abstract
With the expansion of the global mushroom industries, the prediction of future market demand and the production data is important for the further sales of mushroom. The mushroom industries usually receive dynamic demands which are highly non-seasonal and non-periodic in nature. As a result, it is a challenging task to ascertain future mushroom demand and production optimally. In a traditional approach, people produce a certain amount of mushroom in every season based on previous experiences that do not reflect the actual market demand. Therefore, in the case of the shortage of supply, most of the mushroom farms import the products from nearby farms or abroad. Alternatively, the surplus of products than demand is sold to the market at a cheaper rate before the products perish.This paper proposes a machine learning-based solution for the dynamic demand problem in mushroom farms. We have summarized the results obtained from three different machine learning models that are trained with the actual demands of the previous year's mushroom data. After that, we compare the test results given by each of the models to predict the future demand of mushrooms.
Md. Al Maruf, Akramul Azim, Sourojit Mukherjee
ISNCC1
2020 Resource efficient allocation of fog nodes for faster vehicular OTA updates
abstract
Despite reduced network latency and resilience, fog computing has not been leveraged for vehicular Over-the-Air (OTA) updates. Due to vehicle mobility and traffic, the resource utilization of fog nodes is almost non-deterministic, which increases the delay in communication and handover. In this paper, we propose an approach for distributing fog nodes by analyzing the vehicular traffic pattern in a region. The proposed method: (a) finds the optimal number of fog nodes for a specific time interval based on the traffic pattern of a region and (b) maximizes the net reserve resources enabling specific fog nodes. To do so, we use the k-means algorithm to identify traffic load and distribute the fog nodes using our proposed algorithm to maximize fog resource utilization. We present a case study of OTA updates that considers vehicle mobility, data transmission rate, propagation delay and handover delay to predict the required update time. The experimental results demonstrate that the proposed method of fog node allocation extends the net reserve resources by 30.92% on an average, and reduces the OTA update time.
Md. Al Maruf, Anil Singh, Akramul Azim, Nitin Auluck
ISNCC1
2018 Software-based Monitoring for Calibration of Measurement Units in Real-time Systems
abstract
In real-time systems, every task is characterized by its deadline where each task is expected to perform a function producing a correct result within a specified amount of time. A hard real-time system can lead to catastrophic failure if any task misses delivering the correct value at the right time. Although it is very important, most research works in real-time systems avoid discussion on the correctness of values at different points in time. Measurement units or instruments can be integrated with real-time systems to perform sensitive measurements where the measurement accuracy of a device is an essential factor for the precise result. Periodic inspections and calibrations of the measurement units validate the consistent measurement accuracy to ensure the safety of a system. In this paper, we present a software-based monitoring approach for the auto-calibration process that compares sporadically the accuracy of measurement units with the set of determined measurement standards such as National Institute of Standards and Technology (NIST) to ensure the correctness of the measurement instruments. This approach will automatically guide us to correct the measurement errors if the electronic devices are unable to perform with expected accuracy. To explain the applicability of our proposed strategy, we define different techniques considering the availability of the calibration standards and finally show an experiment of anomaly detection in a resistive voltage divider as a case study.
Md. Al Maruf, Akramul Azim
IECON1