Muhammad Shafiq 0003

dblp:27/173-3 · DBLP profile ↗
← Back
27ranked-venue papers
10as first author
16since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 10 · 3 first-author · 7 since 2021Systems, architecture and hardware · 7 · 5 first-author · 1 since 2021Security and privacy · 6 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Building Privacy-Preserving Medical Text Models With a Pretrained Transformer
abstract
The rapid advancement of big data and artificial intelligence (AI) in healthcare heightens the urgency for accurate medical text sentiment analysis. The privacy protection of medical data has been a crucial concern due to its sensitivity. The Internet of Medical Things (IoMT) facilitates large-scale data collection at lower cost, enabling precision medicine. However, decentralized IoMT poses novel challenges to centralized standard encryption schemes. In this article, we propose a novel approach to building privacy-preserving sentiment models with a generative pretrained transformer (GPT). We first convert sensitive medical text data into noise-like and distributed one-hot images. Then, we introduce visual cryptography (VC) for lightweight and secure transmission of medical text across public networks in resource-limited IoMT devices. We adopt a cross-domain sentiment analysis framework that finetunes transformer-based language models for accurate sentiment analysis instead of training GPT in sentiment analysis from scratch. Experimental results show that the proposed approach improves the accuracy and effectiveness of sentiment analysis while maintaining privacy, thereby addressing a significant gap in biomedical text analysis.
Muhammad Shafiq 0003, Lijing Ren, Gautam Srivastava 0001, Denghui Zhang 0001, Sami Bourouis, G. Thippa Reddy
IEEE Internet Things J.1
2024 Optimizing Mobility-Aware Task Offloading in Smart Healthcare for Internet of Medical Things Through Multiagent Reinforcement Learning
abstract
In the scenario of smart healthcare applications, the Internet of Medical Things (IoMT) devices, equipped with limited resources, would offload numerous computation-heavy tasks to an edge server through 5G networks. However, IoMT devices should usually move around different diagnostic areas in smart healthcare systems, leading to the dynamics of the uplink channel quality. Moreover, the burst generation of a substantial number of tasks from IoMT devices can result in congestion within the computing queue of the edge server. And, heterogeneous services in IoMT devices make it hard to collect global information for a central controller to get the optimal optimization for all IoMT devices. So, how to determine task offloading among IoMT devices in a distributed scenario of smart healthcare applications should be considered appropriately and comprehensively. In this paper, we investigate task offloading in mobile edge computing (MEC) through wireless networks. To improve the utilization of wireless resources, non-orthogonal multiple access (NOMA) is adopted in 5G networks. We first formulate the mobility of IoMT devices as a Hidden Markov Model (HMM) and the problem of task offloading policy as a distributed Partial Markov Decision Process (Dec-POMDP). Then, we propose a mobility-aware method based on Multi-agent reinforcement learning for task offloading in 5G NOMA-enabled networks. In our approach, task offloading scheduling for each IoMT device in NOMA-enabled 5G networks is considered to improve energy efficiency and guarantee service quality. Besides, the time complexity and the existence of a Nash equilibrium for our proposed Dec-POMDP method are theoretically derived. Simulations are conducted to show that our algorithm outperforms other alternative methods in energy consumption under the delay constraint.
Chongwu Dong, Yanbin Sun, Muhammad Shafiq 0003, Yuan Liu 0002, Zhihong Tian 0001
IEEE Internet Things J.3
2024 STBCIoT: Securing the Transmission of Biometric Images in Customer IoT
abstract
The recent advancement of the Internet of Things (IoT) and information technology has led to the rapid expansion of interconnectivity among a billion devices across various applications. The advent of massive data has resulted in greater computational dependence, posing obstacles to applying security policies in energy-sensitive devices. However, public-key-based encryption algorithms are impractical or impossible to execute on these resource-limited terminals. In this paper, we propose a lightweight framework called STBCIoT based on a visual cryptography (VC) scheme to achieve low-latency encryption for large-scale data like biometric images. To reduce noise in encryption, we utilize central recognition and gray-level features of QR codes to integrate the visually friendly feature of QR into VC. we further propose a high-quality image generation model with the halftoning effect of VC to improve the quality of decrypted images. The experimental results demonstrate that our proposed method achieves high recognition performance on lossy decrypted images, effectively overcoming the performance limitations of traditional public key encryption methods for largescale images.
Denghui Zhang 0001, Muhammad Shafiq 0003, Gautam Srivastava 0001, G. Thippa Reddy, Le Wang 0008, Zhaoquan Gu
IEEE Internet Things J.2
2024 Multi-Resolution Wavelet Fractal Analysis and Subtask Training for Enhancing Few-Shot Noisy Brainwave Recognition
abstract
The integration of healthcare monitoring with Internet of Things (IoT) networks radically transforms the management and monitoring of human well-being. Portable and lightweight electroencephalography (EEG) systems with fewer electrodes have improved convenience and flexibility while retaining adequate accuracy. However, challenges emerge when dealing with real-time EEG data from IoT devices due to the presence of noisy samples, which impedes improvements in brainwave detection accuracy. Moreover, high inter-subject variability and substantial variability in EEG signals present difficulties for conventional data augmentation and subtask learning techniques, leading to poor generalizability. To address these issues, we present a novel framework for enhancing EEG-based recognition through multi-resolution data analysis, capturing features at different scales using wavelet fractals. The original data can be expanded many times after continuous wavelet transform (CWT) and recombination, alleviating insufficient training samples. In the transfer stage of deep learning (DL) models, we adopt a subtask learning approach to train the recognition model to generalize efficiently. This incorporates wavelets at various scales instead of exclusively considering average prediction performance across scales and paradigms. Through extensive experiments, we demonstrate that our proposed DL-based method excels at extracting features from small-scale and noisy EEG data. This significantly improves healthcare monitoring performance by mitigating the impact of noise introduced by the external environment.
Denghui Zhang 0001, Muhammad Shafiq 0003, Keke Tang, Usman Naseem
IEEE J. Biomed. Health Informatics2
2023 E2EGI: End-to-End Gradient Inversion in Federated Learning
abstract
A plethora of healthcare data is produced every day due to the proliferation of prominent technologies such as Internet of Medical Things (IoMT). Digital-driven smart devices like wearable watches, wristbands and bracelets are utilized extensively in modern healthcare applications. Mining valuable information from the data distributed at the owners' level is useful, but it is challenging to preserve data privacy. Federated learning (FL) has swiftly surged in popularity due to its efficacy in dealing privacy vulnerabilities. Recent studies have demonstrated that Gradient Inversion Attack (GIA) can reconstruct the input data by leaked gradients, previous work demonstrated the achievement of GIA in very limited scenarios, such as the label repetition rate of the target sample being low and batch sizes being smaller than 48. In this paper, a novel method of End-to-End Gradient Inversion (E2EGI) is proposed. Compared to the state-of-the-art method, E2EGI's Minimum Loss Combinatorial Optimization (MLCO) has the ability to realize reconstructed samples with higher similarity, and the Distributed Gradient Inversion algorithm can implement GIA with batch sizes of 8 to 256 on deep network models (such as ResNet-50) and ImageNet datasets. A new Label Reconstruction algorithm is developed that relies only on the gradient information of the target model, which can achieve a label reconstruction accuracy of 81% in one batch sample with a label repetition rate of 96%, a 27% improvement over the state-of-the-art method. This proposed work can underpin data security assessments for healthcare federated learning.
Le Wang 0008, Muhammad Shafiq 0003, Zhaoquan Gu
IEEE J. Biomed. Health Informatics5
2023 FLPK-BiSeNet: Federated Learning Based on Priori Knowledge and Bilateral Segmentation Network for Image Edge Extraction
abstract
Federated learning can effectively ensure data security and improve the problem of data islanding. However, the performance of federated learning-based schemes could be better due to the imbalance of image data. Therefore, this paper proposes a federated learning approach based on priori knowledge and a bilateral segmentation network for image edge extraction. First, federated learning can distribute training images for some special complex images due to the small sample and unshared data. Then, the image with similar edge information to the original image is learned to obtain prior knowledge, and the local uniform sparsity method is used to strengthen the detail features and weaken the background features. Based on the bilateral segmentation network, we introduce a dilated pyramid pooling layer and multi-scale feature fusion module to fuse the shallow detailed features in the context path with the deep abstract features obtained through the dilated pyramid pooling. The final result is obtained by fusing the result with prior knowledge and the result with the context path. Finally, we conduct experiments on some public datasets, and the results show that the proposed method greatly improves extraction accuracy compared with the traditional and the most advanced methods.
Yulong Qiao, Muhammad Shafiq 0003, Gautam Srivastava 0001, Abdul Rehman Javed, G. Thippa Reddy, Shoulin Yin
IEEE Trans. Netw. Serv. Manag.3
2022 IMG-forensics: Multimedia-enabled information hiding investigation using convolutional neural network
abstract
Abstract Information hiding aims to embed a crucial amount of confidential data records into the multimedia, such as text, audio, static and dynamic image, and video. Image‐based information hiding has been a significantly important topic for digital forensics. Here, active image deep steganographic approaches have come forward for hiding data. The least significant bit (LSB) steganography approach is proposed to conceal a secret message into the original image. First, the lightweight stream encryption cryptography encrypts secret information in the cover image to protect embedded information from source to destination. Whereas the encrypted embedded cover information into the carrier of stego‐image with the help of the LSB and then transmit. In the proposed investigational scheme, a convolutional neural net is used. A model is trained to detect and extract patterns of image hidden features, encrypted stego‐image optimization, and classify original and cover images of steganography. Through the experiment result on the forensic image database for mobile steganography of the Center for Statistics and Application in Forensic Evidence, the overall embedded and extracting that the proposed scheme can achieve information hiding as well as revealing with an accuracy rate of 95.1%. The experimental result shows the robustness of the model in terms of efficiency as compared to other state‐of‐the‐art schemes.
Abdullah Ayub Khan, Aftab Ahmed Shaikh, Omar Cheikhrouhou, Asif Ali Laghari, Mamoon Rashid 0001, Muhammad Shafiq 0003, Habib Hamam
IET Image Process.6
2022 A New V-Net Convolutional Neural Network Based on Four-Dimensional Hyperchaotic System for Medical Image Encryption
abstract
In the transmission of medical images, if the image is not processed, it is very likely to leak data and personal privacy, resulting in unpredictable consequences. Traditional encryption algorithms have limited ability to deal with complex data. The chaotic system is characterized by randomness and ergodicity, which has advantages over traditional encryption algorithms in image encryption processing. A novel V-net convolutional neural network (CNN) based on four-dimensional hyperchaotic system for medical image encryption is presented in this study. Firstly, the plaintext medical images are processed into 4D hyperchaotic sequence images, including image segmentation, chaotic system processing, and pseudorandom sequence generation. Then, V-net CNN is used to train chaotic sequences to eliminate the periodicity of chaotic sequences. Finally, the chaotic sequence image is diffused to change the raw image pixel to realize the encryption processing. Simulation test analysis demonstrates that the proposed algorithm has better effect, robustness, and plaintext sensitivity.
Shoulin Yin, Muhammad Shafiq 0003, Asif Ali Laghari, Shahid Karim, Omar Cheikhrouhou, Wajdi Alhakami, Habib Hamam
Secur. Commun. Networks3
2021 Constriction Factor Particle Swarm Optimization based load balancing and cell association for 5G heterogeneous networks
Mohammad Kamrul Hasan 0002, Teong Chee Chuah, Ayman A. El-Saleh, Muhammad Shafiq 0003, Shoaib Ahmed Shaikh, Shayla Islam, Moez Krichen
Comput. Commun.4
2021 DIDDOS: An approach for detection and identification of Distributed Denial of Service (DDoS) cyberattacks using Gated Recurrent Units (GRU)
Mubashir Khaliq, Syed Ibrahim Imtiaz, Aamir Rasool, Muhammad Shafiq 0003, Abdul Rehman Javed, Zunera Jalil, Ali Kashif Bashir
Future Gener. Comput. Syst.5
2021 CorrAUC: A Malicious Bot-IoT Traffic Detection Method in IoT Network Using Machine-Learning Techniques
abstract
Identification of anomaly and malicious traffic in the Internet-of-Things (IoT) network is essential for the IoT security to keep eyes and block unwanted traffic flows in the IoT network. For this purpose, numerous machine-learning (ML) technique models are presented by many researchers to block malicious traffic flows in the IoT network. However, due to the inappropriate feature selection, several ML models prone misclassify mostly malicious traffic flows. Nevertheless, the significant problem still needs to be studied more in-depth that is how to select effective features for accurate malicious traffic detection in the IoT network. To address the problem, a new framework model is proposed. First, a novel feature selection metric approach named CorrAUC is proposed, and then based on CorrAUC, a new feature selection algorithm named CorrAUC is developed and designed, which is based on the wrapper technique to filter the features accurately and select effective features for the selected ML algorithm by using the area under the curve (AUC) metric. Then, we applied the integrated TOPSIS and Shannon entropy based on a bijective soft set to validate selected features for malicious traffic identification in the IoT network. We evaluate our proposed approach by using the Bot-IoT data set and four different ML algorithms. The experimental results analysis showed that our proposed method is efficient and can achieve >96% results on average.
Muhammad Shafiq 0003, Zhihong Tian 0001, Ali Kashif Bashir, Xiaojiang Du, Mohsen Guizani
IEEE Internet Things J.1
2021 An improved watermarking algorithm for robustness and imperceptibility of data protection in the perception layer of internet of things
Mohammad Kamrul Hasan 0002, Samar Kamil, Muhammad Shafiq 0003, Yuvaraj S., Eswaran Saravana Kumar, Rajiv Vincent, Nazmus S. Nafi
Pattern Recognit. Lett.3
2021 Assessing Security of Software Components for Internet of Things: A Systematic Review and Future Directions
abstract
Software component plays a significant role in the functionality of software systems. Component of software is the existing and reusable parts of a software system that is formerly debugged, confirmed, and practiced. The use of such components in a newly developed software system can save effort, time, and many resources. Due to the practice of using components for new developments, security is one of the major concerns for researchers to tackle. Security of software components can save the software from the harm of illegal access and damages of its contents. Several existing approaches are available to solve the issues of security of components from different perspectives in general while security evaluation is specific. A detailed report of the existing approaches and techniques used for security purposes is needed for the researchers to know about the approaches. In order to tackle this issue, the current research presents a systematic literature review (SLR) of the present approaches used for assessing the security of software components in the literature by practitioners to protect software systems for the Internet of Things (IoT). The study searches the literature in the popular and well-known libraries, filters the relevant literature, organizes the filter papers, and extracts derivations from the selected studies based on different perspectives. The proposed study will benefit practitioners and researchers in support of the report and devise novel algorithms, techniques, and solutions for effective evaluation of the security of software components.
Zitian Liao, Shah Nazir, Habib Ullah Khan, Muhammad Shafiq 0003
Secur. Commun. Networks4
2021 Blockchain-Based Automated System for Identification and Storage of Networks
abstract
Network topology is one of the major factors in defining the behavior of a network. In the present scenario, the demand for network security has increased due to an increase in the possibility of attacks by malicious users. In this paper, a blockchain-based system is suggested for securely discovering and storing networks. Techniques such as cloud-based storage systems are not efficient and are lacking in trust, privacy, security, and data control. The blockchain-based technique suggested in this paper is capable of resolving these challenges. Experiments were performed using Mininet, Cisco Packet Tracer, and Ethereum blockchain with the network inference algorithm. This algorithm is capable of inferring the network topology even when only partial information regarding the network is available. The results obtained clearly show that the network is resistant to malicious users and various external attacks, making the network robust.
Deepak Prashar, Nishant Jha, Muhammad Shafiq 0003, Nazir Ahmad, Mamoon Rashid 0001, Shoeib Amin Banday, Habib Ullah Khan
Secur. Commun. Networks3
2021 StFuzzer: Contribution-Aware Coverage-Guided Fuzzing for Smart Devices
abstract
The root cause of the insecurity for smart devices is the potential vulnerabilities in smart devices. There are many approaches to find the potential bugs in smart devices. Fuzzing is the most effective vulnerability finding technique, especially the coverage-guided fuzzing. The coverage-guided fuzzing identifies the high-quality seeds according to the corresponding code coverage triggered by these seeds. Existing coverage-guided fuzzers consider that the higher the code coverage of seeds, the greater the probability of triggering potential bugs. However, in real-world applications running on smart devices or the operation system of the smart device, the logic of these programs is very complex. Basic blocks of these programs play a different role in the process of application exploration. This observation is ignored by existing seed selection strategies, which reduces the efficiency of bug discovery on smart devices. In this paper, we propose a contribution-aware coverage-guided fuzzing, which estimates the contributions of basic blocks for the process of smart device exploration. According to the control flow of the target on any smart device and the runtime information during the fuzzing process, we propose the static contribution of a basic block and the dynamic contribution built on the execution frequency of each block. The contribution-aware optimization approach does not require any prior knowledge of the target device, which ensures our optimization adapting gray-box fuzzing and white-box fuzzing. We designed and implemented a contribution-aware coverage-guided fuzzer for smart devices, called StFuzzer. We evaluated StFuzzer on four real-world applications that are often applied on smart devices to demonstrate the efficiency of our contribution-aware optimization. The result of our trials shows that the contribution-aware approach significantly improves the capability of bug discovery and obtains better execution speed than state-of-the-art fuzzers.
Jiageng Yang, Xinguo Zhang, Hui Lu 0005, Muhammad Shafiq 0003, Zhihong Tian 0001
Secur. Commun. Networks4
2021 Communication Delay Modeling for Wide Area Measurement System in Smart Grid Internet of Things Networks
abstract
We present communication frameworks, models, and protocols of smart grid Internet of Things (IoT) networks based on the IEEE and IEC standards. The measurement, control, and monitoring of grid being achieved through phasor measurement unit (PMU) based wide area measurement (WAM) framework. The WAM framework applied the IEEE standard C37.118 phasor exchange protocol to collect grid data from various substation devices. The existing frameworks include the IEC 61850 protocol and programmable logic controllers (PLCs) based supervisory control and data acquisition (SCADA) system. These protocols have been selected as per the smart grid configuration and communication design. However, the existing frameworks have severe synchronization errors due to the communication delays of IoT networks in the smart grid. Therefore, this article designs the timing mechanism and a delay model to reduce the timing delay and boost real‐time measurement, monitoring, and control performance of the smart grid WAM applications. The result shows that the proposed model outperformed the existing WAM system.
Mohammad Kamrul Hasan 0002, Shayla Islam, Muhammad Shafiq 0003, Fatima Rayan Awad Ahmed, Somya Khidir Mohmmed Ataelmanan, Nissrein Babiker Mohammed Babiker, Khairul Azmi Abu Bakar
Wirel. Commun. Mob. Comput.3
2020 IoT malicious traffic identification using wrapper-based feature selection mechanisms
abstract
Machine Learning (ML) plays very significant role in the Internet of Things (IoT) cybersecurity for malicious and intrusion traffic identification. In other words, ML algorithms are widely applied for IoT traffic identification in IoT risk management . However, due to inaccurate feature selection, ML techniques misclassify a number of malicious traffic in smart IoT network for secured smart applications. To address the problem, it is very important to select features set that carry enough information for accurate smart IoT anomaly and intrusion traffic identification. In this paper, we firstly applied bijective soft set for effective feature selection to select effective features, and then we proposed a novel CorrACC feature selection metric approach. Afterward, we designed and developed a new feature selection algorithm named Corracc based on CorrACC, which is based on wrapper technique to filter the features and select effective feature for a particular ML classifier by using ACC metric. For the evaluation our proposed approaches, we used four different ML classifiers on the BoT-IoT dataset. Experimental results obtained by our algorithms are promising and can achieve more than 95% accuracy.
Muhammad Shafiq 0003, Zhihong Tian 0001, Ali Kashif Bashir, Xiaojiang Du, Mohsen Guizani
Comput. Secur.1
2020 Selection of effective machine learning algorithm and Bot-IoT attacks traffic identification for internet of things in smart city
Muhammad Shafiq 0003, Zhihong Tian 0001, Yanbin Sun, Xiaojiang Du, Mohsen Guizani
Future Gener. Comput. Syst.1
2020 SoftSystem: Smart Edge Computing Device Selection Method for IoT Based on Soft Set Technique
abstract
The Internet of Things (IoT) is growing day by day, and new IoT devices are introduced and interconnected. Due to this rapid growth, IoT faces several issues related to communication in the edge computing network. The critical issue in these networks is the effective edge computing IoT device selection whenever there are several edge nodes to carry information. To overcome this problem, in this paper, we proposed a new framework model named SoftSystem based on the soft set technique that recommends useful IIoT devices. Then, we proposed an algorithm named Softsystemalgo. For the proposed system, three different parameters are selected: IoT Device Security (IDSC), IoT Device Storage (IDST), and IoT Device Communication Speed (IDCS). We also find out the most significant parameters from the given set of parameters. It is evident that our proposed system is effective for the selection of edge computing devices in the IoT network.
Muhammad Shafiq 0003, Zhihong Tian 0001, Ali Kashif Bashir, Korhan Cengiz, Adnan Tahir
Wirel. Commun. Mob. Comput.1
2020 An adaptive heuristic for managing energy consumption and overloaded hosts in a cloud data center
Weizhe Zhang, Keqin Li 0001, Chuanyi Liu, Muhammad Shafiq 0003, Nabin Kumar Karn
Wirel. Networks5
2018 WeChat traffic classification using machine learning algorithms and comparative analysis of datasets
abstract
In this research paper, we present the first classification study to classify WeChat application service flow traffic (text messages, picture messages, audio call and video call traffic) classification and secondly to find out the effectiveness of big dataset and small dataset as well as to find out effective machine learning classifiers. We firstly capture WeChat traffic and then extract 44 features then we combine capture traffic to make full instance of dataset. Then we make reduce instances of dataset from the full instance of dataset to show the effectiveness of large dataset and small dataset. Then we execute well known machine learning classifiers. Using statistical test, we use Wilcoxon and Friedman statistical test for the datasets and ML classifiers to find more deeply its effectiveness. Experimental results show that reduce instance dataset show high accuracy result compared to full instance and C4.5 classifier perform effectively as compared to other classifiers.
Muhammad Shafiq 0003, Xiangzhan Yu, Asif Ali Laghari
Int. J. Inf. Comput. Secur.1
2018 A machine learning approach for feature selection traffic classification using security analysis
Muhammad Shafiq 0003, Xiangzhan Yu, Ali Kashif Bashir, Hassan Nazeer Chaudhry
J. Supercomput.1
2018 An Efficient Security System for Mobile Data Monitoring
abstract
During the last decade, rapid development of mobile devices and applications has produced a large number of mobile data which hide numerous cyber‐attacks. To monitor the mobile data and detect the attacks, NIDS/NIPS plays important role for ISP and enterprise, but now it still faces two challenges, high performance for super large patterns and detection of the latest attacks. High performance is dominated by Deep Packet Inspection (DPI) mechanism, which is the core of security devices. A new TTL attack is just put forward to escape detecting, such that the adversary inserts packet with short TTL to escape from NIDS/NIPS. To address the above‐mentioned problems, in this paper, we design a security system to handle the two aspects. For efficient DPI, a new two‐step partition of pattern set is demonstrated and discussed, which includes first set‐partition and second set‐partition. For resisting TTL attacks, we set reasonable TTL threshold and patch TCP protocol stack to detect the attack. Compared with recent produced algorithm, our experiments show better performance and the throughput increased 27% when the number of patterns is 106. Moreover, the success rate of detection is 100%, and while attack intensity increased, the throughput decreased.
Likun Liu, Hongli Zhang 0001, Xiangzhan Yu, Yi Xin 0002, Muhammad Shafiq 0003, Mengmeng Ge 0003
Wirel. Commun. Mob. Comput.5
2013 A template system for the efficient compilation of domain abstractions onto reconfigurable computers
Muhammad Shafiq 0003, Miquel Pericàs, Nacho Navarro, Eduard Ayguadé
J. Syst. Archit.1
2011 Assessing Accelerator-Based HPC Reverse Time Migration
abstract
Oil and gas companies trust Reverse Time Migration (RTM), the most advanced seismic imaging technique, with crucial decisions on drilling investments. The economic value of the oil reserves that require RTM to be localized is in the order of 10^{13} dollars. But RTM requires vast computational power, which somewhat hindered its practical success. Although, accelerator-based architectures deliver enormous computational power, little attention has been devoted to assess the RTM implementations effort. The aim of this paper is to identify the major limitations imposed by different accelerators during RTM implementations, and potential bottlenecks regarding architecture features. Moreover, we suggest a wish list, that from our experience, should be included as features in the next generation of accelerators, to cope with the requirements of applications like RTM. We present an RTM algorithm mapping to the IBM Cell/B.E., NVIDIA Tesla and an FPGA platform modeled after the Convey HC-1. All three implementations outperform a traditional processor (Intel Harpertown) in terms of performance (10x), but at the cost of huge development effort, mainly due to immature development frameworks and lack of well-suited programming models. These results show that accelerators are well positioned platforms for this kind of workload. Due to the fact that our RTM implementation is based on an explicit high order finite difference scheme, some of the conclusions of this work can be extrapolated to applications with similar numerical scheme, for instance, magneto-hydrodynamics or atmospheric flow simulations.
Mauricio Araya-Polo, Javier Cabezas, Mauricio Hanzich, Miquel Pericàs, Félix Rubio, Isaac Gelado, Muhammad Shafiq 0003, Enric Morancho, Nacho Navarro, Eduard Ayguadé, José María Cela, Mateo Valero
IEEE Trans. Parallel Distributed Syst.7
2010 FEM: A Step Towards a Common Memory Layout for FPGA Based Accelerators
abstract
FPGA devices are mostly utilized for customized application designs with heavily pipelined and aggressively parallel computations. However, little focus is normally given to the FPGA memory organizations to efficiently use the data fetched into the FPGA. This work presents a Front End Memory (FEM) layout based on BRAMs and Distributed RAM for FPGA-based accelerators. The presented memory layout serves as a template for various data organizations which is in fact a step towards the standardization of a methodology for FPGA based memory management inside an accelerator. We present example application kernels implemented as specializations of the template memory layout. Further, the presented layout can be used for Spatially Mapped-Shared Memory multi-kernel applications targeting FPGAs. This fact is evaluated by mapping two applications, an Acoustic Wave Equation code and an N-Body method, to three multi-kernel execution models on a Virtex-4 L×200 device. The results show that the shared memory model for Acoustic Wave Equation code outperforms the local and runtime reconfigured models by 1.3-1.5×, respectively. For the N-Body method the shared model is slightly more efficient with a small number of bodies, but for larger systems the runtime reconfigured model shows a 3× speedup over the other two models.
Muhammad Shafiq 0003, Miquel Pericàs, Nacho Navarro, Eduard Ayguadé
FPL1
2009 Exploiting memory customization in FPGA for 3D stencil computations
abstract
3D stencil computations are compute-intensive kernels often appearing in high-performance scientific and engineering applications. The key to efficiency in these memory-bound kernels is full exploitation of data reuse. This paper explores the design aspects for 3D-Stencil implementations that maximize the reuse of all input data on a FPGA architecture. The work focuses on the architectural design of 3D stencils with the form n × (n + 1) × n, where n = {2, 4, 6, 8, ...}. The performance of the architecture is evaluated using two design approaches, ¿Multi-Volume¿ and ¿Single-Volume¿. When n = 8, the designs achieve a sustained throughput of 55.5 GFLOPS in the ¿Single-Volume¿ approach and 103 GFLOPS in the ¿Multi-Volume¿ design approach in a 100-200 MHz multi-rate implementation on a Virtex-4 LX200 FPGA. This corresponds to a stencil data delivery of 1500 bytes/cycle and 2800 bytes/cycle respectively. The implementation is analyzed and compared to two CPU cache approaches and to the statically scheduled local stores on the IBM PowerXCell 8i. The FPGA approaches designed here achieve much higher bandwidth despite the FPGA device being the least recent of the chips considered. These numbers show how a custom memory organization can provide large data throughput when implementing 3D stencil kernels.
Muhammad Shafiq 0003, Miquel Pericàs, Raúl de la Cruz, Mauricio Araya-Polo, Nacho Navarro, Eduard Ayguadé
FPT1