EDBT 2026 Demo / reviewers in the wild / expert
Syed Rafay Hasan
dblp:69/1219
· DBLP profile ↗
33ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0003-0183-8086ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 30 · 4 first-author · 7 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mitigation of Camouflaged Adversarial Attacks in Autonomous Vehicles-A Case Study Using CARLA SimulatorabstractAutonomous vehicles (AVs) rely heavily on cameras and artificial intelligence (AI) to make safe and accurate driving decisions. However, since AI is the core enabling technology, this raises serious cyber threats that hinder the large-scale adoption of AVs. Therefore, it is crucial to analyze the resilience of AV security systems against sophisticated attacks that manipulate camera inputs and deceive AI models. In this paper, we develop camera-camouflaged adversarial attacks and asses their impact on traffic sign recognition (TSR) in AVs. Specifically, the camera-camouflaged attacks are initiated by modifying the texture of a stop sign to fool the AV’s object detection system, thereby affecting the AV actuators. The attacks’ effectiveness is tested using the CARLA AV simulator, and the results show that such attacks can delay the auto-braking response to the stop sign, resulting in potentially catastrophic situations. We conduct extensive experiments under various conditions and on different CARLA maps, confirming that the proposed attacks are effective and robust. Then, two defense strategies are presented to mitigate the effect of camera-camouflaged attacks on the stop sign recognition. The proposed attack and defense methods are applicable to other end-to-end trained autonomous cyber-physical systems. Yago Romano Martinez, Brady Carter, Abhijeet Solanki, Wesam Al Amiri, Syed Rafay Hasan, Nan Guo 0001 |
ISCAS | 5 |
| 2025 | SHEATH: Defending Horizontal Collaboration for Distributed CNNs Against Adversarial NoiseabstractAs edge computing and the Internet of Things (IoT) expand, horizontal collaboration (HC) emerges as a distributed data processing solution for resource-constrained devices. In particular, a convolutional neural network (CNN) model can be deployed on multiple IoT devices, allowing distributed inference execution for image recognition while ensuring model and data privacy. Yet, this distributed architecture remains vulnerable to adversaries who want to make subtle alterations that impact the model, even if they lack access to the entire model. Such vulnerabilities can have severe implications for various sectors, including healthcare, military, and autonomous systems. However, security solutions for these vulnerabilities have not been explored. This paper presents a novel framework for Secure Horizontal Edge with Adversarial Threat Handling (SHEATH) to detect adversarial noise and eliminate its effect on CNN inference by recovering the original feature maps. Specifically, SHEATH aims to address vulnerabilities without requiring complete knowledge of the CNN model in HC edge architectures based on sequential partitioning. It ensures data and model integrity, offering security against adversarial noise in diverse HC environments. Our evaluations demonstrate SHEATH’s adaptability and effectiveness across diverse CNN configurations. Muneeba Asif, Mohammad Kumail Kazmi, Mohammad Ashiqur Rahman, Syed Rafay Hasan, Soamar Homsi |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | System Integration of Xilinx DPU and HDMI for Real-Time Inference in PYNQ Environment With Image EnhancementabstractUse of edge computing in application of Computer Vision (CV) is an active field of research. Today, most CV applications make use of Convolutional Neural Networks (CNNs) to inference on and interpret video data. These edge devices are responsible for several CV related tasks, such as gathering, processing and enhancing, inferencing on, and displaying video data. Due to ease of reconfiguration, computation on FPGA fabric is used to achieve such complex computation tasks. Xilinx provides the PYNQ environment as a user-friendly interface that facilitates in Hardware/Software system integration. However, to the best of authors’ knowledge there is no end-to-end framework available for the PYNQ environment that allows Hardware/Software system integration and deployment of CNNs for real-time input feed from High Definition Multimedia Interface (HDMI) input to HDMI output, along with insertion of customized hardware IPs. In this work we propose an integration of reaL-time image Enancement IP with AI inferencing engine in the Pynq environment (LEAP), that integrates HDMI, AI acceleration, image enhancement in the PYNQ environment for Xilinx’s Microprocessor on Chip (MPSoC) platform. We evaluate our methodology with two well known CNN models, Resnet50 and YOLOv3. To validate our proposed methodology, LEAP, a simple image enhancement algorithm, histogram equalization, is designed and integrated in the FPGA fabric along with Xilinx’s Deep Processing Unit (DPU). Our results show successful implementation of end-to-end integration using completely open source information. Jonathan J. Sanderson, Syed Rafay Hasan |
ISCAS | 2 |
| 2024 | Detection and Mitigation of Subtle Feature-map Attacks in Pseudo Parallel Collaborative CNN Models for Distributed Edge IntelligenceabstractAlthough Collaborative Deep Neural Network (CDNN) promises to be an alternative mechanism to mitigate the effects of the untrusted cloud, this approach is susceptible to other kinds of adversarial attacks, which arise from one or more untrusted devices in CDNN acting maliciously. However, since each untrusted node in the CDNN contains only partial information of the complete DNN model, it is worth investigating whether the attacker can still muster a viable threat to CDNN or not. This led to the investigation of attack scenarios and their effects on convolutional neural networks (CNN) used for image classification in the CDNN environment. In this research, we are investigating the shortcomings of existing attacks on CDNN, that lead to non-subtle attack to the defender who is on look out against such attacks. Our research showed that sparse nature of feature maps (FMs) due to the ReLU function lead to many existing attacks more obvious to the attacker. Next, we investigated how one can detect the existing attacks if the defender has some previous knowledge of the complete CNN’s FMs. Our results show minimal detection overhead of about 2%, with an accuracy of 95% and F1 score of above 0.97.1 Syed Rafay Hasan, Mohammad Ashiqur Rahman, Soamar Homsi |
VTC Fall | 2 |
| 2023 | Enhancing the Security of Collaborative Deep Neural Networks: An Examination of the Effect of Low Pass FiltersabstractTo ensure that accuracy and latency are not compromised while deploying Deep Neural Networks (DNNs) on edge devices, trained DNN models can be partitioned across many collaborating edge devices for inference. However, this collaborative inference paradigm raises new security risks because one of the collaborating edge devices could be malicious or compromised, leading to compromised accuracy and reliability of inference results. To address this challenge, this paper explores the use of low-pass filters to enhance the robustness of Collaborative DNNs. The study deploys a VGG16 network, trained on the German Traffic Sign Recognition Benchmarks (GTSRB) dataset, and a MobileNet network trained on the ImageNet dataset, using two prevalent collaborative inference methodologies. The output feature maps (FMs) of a vulnerable edge device are perturbed using four advanced adversarial noises, namely Speckle, Salt-and-Pepper, Gaussian noise, and the Fast Gradient Signed Method (FGSM). Experimental results demonstrate that implementing low-pass filtering can significantly enhance the robustness of Collaborative DNNs. On average, the top-1 classification accuracy is improved by 2.1x times, making the DNNs more robust to adversarial attacks. Adewale Adeyemo, Syed Rafay Hasan |
ACM Great Lakes Symposium on VLSI | 2 |
| 2022 | LaBaNI: Layer-based Noise Injection Attack on Convolutional Neural NetworksabstractHardware accelerator-based CNN inference improves the performance and latency but increases the time-to-market. As a result, CNN deployment on hardware is often outsourced to untrusted third parties (3Ps) with security risks, like hardware Trojans (HTs). Therefore, during the outsourcing, designers conceal the information about initial and final CNN layers from 3Ps. However, this paper shows that this solution is ineffective by proposing a hardware-intrinsic attack (HIA), Layer-based Noise Injection (LaBaNI), which successfully performs misclassification without knowing the initial and final layers. LaBaNi uses the statistical properties of feature maps of the CNN to design the trigger with a very low triggering probability and a payload for misclassification. To show the effectiveness of LaBaNI, we demonstrated it on LeNet and LeNet-3D CNN models deployed on Xilinx's PYNQ board. In the experimental results, the attack is successful, non-periodic, and random, hence difficult to detect. Results show that LaBaNI utilizes up to 4% extra LUTs, 5% extra DSPs, and 2% extra FFs, respectively. Tolulope A. Odetola, Faiq Khalid, Syed Rafay Hasan |
ACM Great Lakes Symposium on VLSI | 3 |
| 2022 | Towards Enabling Dynamic Convolution Neural Network Inference for Edge IntelligenceabstractDeep learning applications have achieved great success in numerous real-world applications. Deep learning models, especially Convolution Neural Networks (CNN) are often prototyped using FPGA because it offers high power efficiency, and reconfigurability. The deployment of CNNs on FPGAs follows a design cycle that requires saving of model parameters in the on-chip memory during High level synthesis (HLS). Recent advances in edge intelligence requires CNN inference on edge network to increase throughput and reduce latency. To provide flexibility, dynamic parameter allocation to different mobile devices is required to implement either a predefined or defined on-the-fly CNN architecture. In this study, we present novel methodologies for dynamically streaming the model parameters at run-time to implement a traditional CNN architecture. We further propose a library-based approach to design scalable and dynamic distributed CNN inference on the fly leveraging partial-reconfiguration techniques, which is particularly suitable for resource constrained edge devices. The proposed techniques are implemented on the Xilinx PYNQ-Z2 board to prove the concept by utilizing the LeNet-5 CNN model. The results show that the proposed methodologies are effective, with classification accuracy rates of 92%, 86% and 94% respectively. Adewale Adeyemo, Travis Sandefur, Tolulope A. Odetola, Syed Rafay Hasan |
ISCAS | 4 |
| 2022 | Framework to Benchmark CNNs (FaBCNN) for Processing Real-Time HD Video Streams on FPGAsabstractThe deployment of Convolutional Neural Networks (CNNs) on resource-constrained edge devices for inference is challenging due to its computation, memory, energy, and bandwidth requirements. To address these issues, FPGAs are commonly used to implement CNNs because of their high flexibility and low power consumption. This paper proposes a methodology that provides a technique to benchmark CNNs using HDMI input and output in real-time with 720p high definition (HD) resolution. This methodology can be utilized in a classroom set up to teach CNN and computer vision fundamentals. To illustrate the effectiveness of the proposed methodology, several object detection and image classification CNNs were deployed on the Xilinx ZCU104 FPGA board. Video is provided to the FPGA in real-time from an HDMI input source. The output of a given CNN is converted to an HDMI stream and displayed on a separate monitor at 720p HD resolution. The experimental results show that this methodology can perform object detection and image classification on real-time video at speeds of around 10 FPS and 30 FPS, respectively. Travis Sandefur, Syed Rafay Hasan |
ISCAS | 2 |
| 2022 | Edge Intelligence in Mobile Nodes: Opportunistic Pipeline via 5G D2D for On-site SensingabstractForeseeing several potential use cases, this paper proposes a node-level edge intelligence (EI) based mobile pipeline computing concept in a Device-to-Device (D2D) communication setup and studies the related issues, where D2D is likely based on millimeter-wave (mmWave) in the 5G mobile communication. Different from traditional distributed computing systems, the proposed opportunistic system employs a chain of wirelessly pipelined resource-limited node-level edge devices (NEDs) on the move to handle real-time computation-intensive multi-stage processing for which current cloud computing technology may not be suitable. The feasibility of such a mobile pipeline can be anticipated as high-speed and low-latency wireless technologies get mature. We present a system model by defining the architecture and basic functions and introducing a possible optimal pipeline finding procedure. As part of the feasibility assessment, the impact of mmWave blockage on the pipeline stability is analyzed and examined for both single-pipeline and concurrent-multiple-pipeline scenarios. Our design and analysis results provide specific insight to guide system design and lay a foundation for further work along this line. Nan Guo 0001, Hawzhin Mohammed, Syed Rafay Hasan |
VTC Fall | 3 |
| 2021 | SoWaF: Shuffling of Weights and Feature Maps: A Novel Hardware Intrinsic Attack (HIA) on Convolutional Neural Network (CNN)abstractSecurity of inference phase deployment of Convolutional neural network (CNN) into resource constrained embedded systems (e.g. low end FPGAs) is a growing research area. Using secure practices, third party FPGA designers can be provided with no knowledge of initial and final classification layers. In this work, we demonstrate that hardware intrinsic attack (HIA) in such a "secure" design is still possible. Proposed HIA is inserted inside mathematical operations of individual layers of CNN, which propagates erroneous operations in all the subsequent CNN layers that leads to misclassification. The attack is non-periodic and completely random, hence it becomes difficult to detect. Five different attack scenarios with respect to each CNN layer are designed and evaluated based on the overhead resources and the rate of triggering in comparison to the original implementation. Our results for two CNN architectures show that in all the attack scenarios, additional latency is negligible (<; 0.61%), increment in DSP, LUT, FF is also less than 2.36%. Three attack scenarios does not require any additional BRAM resources, while in two scenarios BRAM increases, which compensates with the corresponding decrease in FF and LUTs. To the authors' best knowledge this work is the first to address the hardware intrinsic CNN attack with attacker does not have knowledge of the full CNN. Tolulope A. Odetola, Syed Rafay Hasan |
ISCAS | 2 |
| 2020 | MacLeR: Machine Learning-Based Runtime Hardware Trojan Detection in Resource-Constrained IoT Edge DevicesabstractTraditional learning-based approaches for runtime hardware Trojan (HT) detection require complex and expensive on-chip data acquisition frameworks, and thus incur high area and power overhead. To address these challenges, we propose to leverage the power correlation between the executing instructions of a microprocessor to establish a machine learning (ML)-based runtime HT detection framework, called MacLeR. To reduce the overhead of data acquisition, we propose a single power-port current acquisition block using current sensors in time-division multiplexing, which increases accuracy while incurring reduced area overhead. We have implemented a practical solution by analyzing multiple HT benchmarks inserted in the RTL of a system-on-chip (SoC) consisting of four LEON3 processors integrated with other IPs, such as vga_lcd, RSA, AES, Ethernet, and memory controllers. Our experimental results show that compared to state-of-the-art HT detection techniques, MacLeR achieves 10% better HT detection accuracy (i.e., 96.256%) while incurring a 7× reduction in area and power overhead (i.e., 0.025% of the area of the SoC and <; 0.07% of the power of the SoC). In addition, we also analyze the impact of process variation (PV) and aging on the extracted power profiles and the HT detection accuracy of MacLeR. Our analysis shows that variations in fine-grained power profiles due to the HTs are significantly higher compared to the variations in fine-grained power profiles caused by the PVs and aging effects. Moreover, our analysis demonstrates that on average, the HT detection accuracy drops in MacLeR is less than 1% and 9% when considering only PV and PV with worst case aging, respectively, which is ≈10× less than in the case of the state-of-the-art ML-based HT detection technique. Faiq Khalid, Syed Rafay Hasan, Sara Zia, Osman Hasan, Falah R. Awwad, Muhammad Shafique 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | MulMapper: Towards an Automated FPGA-Based CNN Processor Generator Based on a Dynamic Design Space ExplorationabstractMany enterprises are adopting deep learning algorithms in their everyday tasks faster than ever. Convolutional Neural Networks (CNNs) in particular are being used widely due to the impressive performance in various application areas. FPGAs, on the other hand, are becoming a promising hardware platform for various deep learning algorithms including CNN. However, optimized and efficient FPGA design requires an expert with hardware design skills. This is particularly a challenge for deep learning practitioners who would like to accelerate their algorithm without worrying about the underlying hardware knowledge required to accomplish that in FPGAs. In this work we are proposing an automated framework, MulMapper, that can generate a functional and synthesized CNN processor hardware IP (using Vivado HLS) for Zynq-based FPGAs, given Caffe-based CNN definition file. We created a dynamic and novel design space utilizing Target Device Resource, Target Core Mode and Target Data Width as design space dimensions. MulMapper explores the design space in these three dimensions and proposes the optimum design points. We tested MulMapper framework on common CNN architectures, LeNet, CNP and CIFAR-10. It has been verified that early-stage MulMapper can lead to synthesis of resource-optimized CNN processor hardware IP that can be used for many regular CNN variants. Comparison with the state-of-the-art shows that architectures generated using MulMapper obtained up to 25-29× DSP48 and 13-20× on-chip memory reduction, with up to 0.35 GOP/sec performance. Muluken Hailesellasie, Syed Rafay Hasan, Otmane Aït Mohamed |
ISCAS | 2 |
| 2019 | Survey on recent counterfeit IC detection techniques and future research directions
Enahoro Oriero, Syed Rafay Hasan |
Integr. | 2 |
| 2018 | FPGA-Based Convolutional Neural Network Architecture with Reduced Parameter RequirementsabstractThe success of deep learning has fast paced the evolution of current technology at unprecedented rate. In particular, deep convolutional neural networks (CNNs) has gained a lot of attention due to their extraordinary performance in a wide range of computer vision applications. While the performance of CNNs has been excellent, their implementation complexity has, however, always posed a challenge due to their computational and memory access intensive nature of CNNs especially for resource constrained embedded platforms. In this paper, we propose a novel reduced-parameter CNN architecture that can be used for image classification applications, which results in a significant network model size reduction. Our reduction method, inspired by SqueezeNet, replaces convolutional layer kernels with smaller sized kernels and removes all the fully connected layers other than the last classifying layer. The proposed architecture results in less computational complexity when deployed in hardware. We implemented the proposed architecture by fitting all trained network parameters on-chip using Xilinx Vivado targeting Zynq XC7Z020-1CLG484C FPGA device. The proposed architecture has 11.2× less parameters and has an improvement of 2.8× Area-Delay Product, compared to LeNet, resulting in an efficient hardware deployment. Muluken Hailesellasie, Syed Rafay Hasan, Faiq Khalid, Falah R. Awwad, Muhammad Shafique 0001 |
ISCAS | 2 |
| 2018 | Low Power Digital Clock Multipliers for Battery-Operated Internet of Things (IoT) DevicesabstractThe recent advancements in system-on-chip (SoC) and network-on-chip (NoC) have enormously increased the number of on-chip frequency domains that are originating from multiple on-chip clock sources. In modern battery-operated internet of things (IoT) devices, limited power budget and requirement for complex clock distribution schemes increases the usage clock multipliers. These multiple clock signal requirements are usually catered for by using frequency multipliers with clock generators. However, most of these multipliers are based on analog components that require a customized layout, involve timing uncertainties, and are power hungry and highly prone to mismatches in the process variations and environmental changes. Moreover, in modern battery-operated smart devices for IoT have very limited power budget, which makes the design of clock multipliers even more challenging. To address these issues, we propose a delay-based digital frequency multiplier, which uses 2-input XNOR gates and a true single-phase clock (TSPC) flip-flop because of pulse generation and edge detection properties, respectively. The proposed multiplier is based on the digital components, therefore, it reduces the power consumption significantly, i.e., 1.6mW, which is almost 50% lesser than other low power state-of-the-art designs. Moreover, it can operate for a wide range of input frequencies, ~400MHz to 1GHz. The Monte-Carlo simulation results are very promising as they indicate the robustness of the design against process and environmental variations. Faiq Khalid, Sunil Nanjiani, Syed Rafay Hasan, Osman Hasan, Falah R. Awwad, Muhammad Shafique 0001 |
ISCAS | 3 |
| 2018 | Runtime hardware Trojan monitors through modeling burst mode communication using formal verification
Faiq Khalid, Syed Rafay Hasan, Osman Hasan, Falah R. Awwad |
Integr. | 2 |
| 2017 | Power profiling of microcontroller's instruction set for runtime hardware Trojans detection without golden circuit modelsabstractGlobalization trends in integrated circuit (IC) design are leading to increased vulnerability of ICs against hardware Trojans (HT). Recently, several side channel parameters based techniques have been developed to detect these hardware Trojans that require golden circuit as a reference model, but due to the widespread usage of IPs, most of the system-on-chip (SoC) do not have a golden reference. Hardware Trojans in intellectual property (IP)-based SoC designs are considered as major concern for future integrated circuits. Most of the state-of-the-art runtime hardware Trojan detection techniques presume that Trojans will lead to anomaly in the SoC integration units. In this paper, we argue that an intelligent intruder may intrude the IP-based SoC without disturbing the normal SoC operation or violating any protocols. To overcome this limitation, we propose a methodology to extract the power profile of the micro-controllers instruction sets, which is in turn used to train a machine learning algorithm. In this technique, the power profile is obtained by extracting the power behavior of the micro-controllers for different assembly language instructions. This trained model is then embedded into the integrated circuits at the SoC integration level, which classifies the power profile during runtime to detect the intrusions. We applied our proposed technique on MC8051 micro-controller in VHDL, obtained the power profile of its instruction set and then applied deep learning, k-NN, decision tree and naive Bayesian based machine learning tools to train the models. The cross validation comparison of these learning algorithm, when applied to MC8051 Trojan benchmarks, shows that we can achieve 87% to 99% accuracy. To the best of our knowledge, this is the first work in which the power profile of a microprocessor's instruction set is used in conjunction with machine learning for runtime HT detection. Faiq Khalid, Syed Rafay Hasan, Osman Hasan, Falah R. Awwad |
DATE | 2 |
| 2017 | A fast FPGA-based deep convolutional neural network using pseudo parallel memoriesabstractDeep learning is gaining popularity in the recent years due to its impressive performance in different application areas. Convolutional Neural Network (CNN) is the state-of-the-art deep learning architecture that is being used widely in the areas of image recognition, speech recognition and many other applications. CNN is computationally intensive and resource hungry architecture. Hence, its efficient hardware implementation is one of the challenges faced by researchers. FPGAs are the dominating platform choice when it comes to implementation of such architectures. This paper presents an efficient implementation of convolutional layer of CNN, that can substantially reduce the long computational time by utilizing the parallel usage of memories. The technique proposed distributes the input image to be classified into P memories; where P is obtained as an optimum trade-off between number of clock cycles and memory resources. Reading concurrently from all P memories reduces the required number of clock cycles proportionally, at the expense of complex control unit. The proposed architecture is implemented using Xilinx Vivado, targeting Zynq XC7Z020-1CLG484C device. The architecture is tested using MNIST dataset and successfully compared with Caffe's convolutional layer output. Muluken Hailesellasie, Syed Rafay Hasan |
ISCAS | 2 |
| 2017 | Motion artifact reduction from PPG signals during intense exercise using filtered X-LMSabstractPhotoplethysomographic (PPG) signal is crucial for non-invasive monitoring of heart rate. It is acquired by using pulse oximeter that are prone to artifacts. A major application of this technique is monitoring the heart rate during physical exertion. Extraction of heart rate (HR) from the PPG in this case is difficult due to the strong motion related artifacts. This paper proposes an efficient method based on a reference generation using singular value decomposition and then multistage application of filtered X-LMS for removing motion artifacts from PPG signal. Simultaneous three-axis acceleration data is acquired and used as reference signal to measure time and extent of motion artifact in PPG signal. This is followed by an application of Slope Sum Method (SSM) to track peaks, and thus determine the heart rate. Testing of proposed method on PPG signals acquired from multiple subjects performing intense exercises (jogging at an average speed of 12 km/hour), results in mean absolute error of 1.37 beats per minute (BPM). Moreover, it is shown that proposed algorithm is robust to excessive occurrence of motion artifacts. Khawaja Taimoor Tanweer, Syed Rafay Hasan, Awais M. Kamboh |
ISCAS | 2 |
| 2017 | Self-triggering hardware trojan: Due to NBTI related aging in 3-D ICs
Siraj Fulum Mossa, Syed Rafay Hasan, Omar S. Elkeelany |
Integr. | 2 |
| 2017 | Hardware trojans in 3-D ICs due to NBTI effects and countermeasure
Siraj Fulum Mossa, Syed Rafay Hasan, Omar S. Elkeelany |
Integr. | 2 |
| 2016 | Synchronously triggered GALS design templates leveraging QDI asynchronous interfacesabstractSingle clock distribution over a large high performance chip can be very challenging. This led to evolution of globally asynchronous and locally Synchronous (GALS) systems in modern deep sub-micron (DSM) technology. In GALS mostly bundled data protocols which are based on handshake mechanism, are used for data transfer. But these protocols rely on timing assumptions between handshake signals and data values that causes timing closure problems, which poses strict constraints in system-on-chip (SoC) design. This work leverages quasi delay insensitive (QDI) designs to propose GALS design templates. This will facilitate the use of GALS systems in a conventional digital design flow with minimal intervention to interfacing modules. Modifications for two different quasi delay insensitive (QDI) asynchronous designs have been suggested, implemented and verified by using the proposed templates. Power, energy and latency have been compared for two different interfaces. Waqas Gul, Syed Rafay Hasan, Osman Hasan, Faiq Khalid, Falah R. Awwad |
ISCAS | 2 |
| 2016 | A self-learning framework to detect the intruded integrated circuitsabstractGlobalization trends in integrated circuit (IC) design using deep submicron (DSM) technologies are leading to increased vulnerability of ICs against malicious intrusions. These malicious intrusions are referred as hardware Trojans. One way to address this threat is to utilize unique electrical signatures of ICs. However, this technique requires analyzing extensive sensor data to detect the intruded integrated circuits. In order to overcome this limitation, we propose to combine the signature extraction mechanism with machine learning algorithms to develop a self-learning framework that can detect the intruded integrated circuits. The proposed approach applies the lazy, eager or probabilistic learners to generate self-learning prediction model based on the electrical signatures. In order to validate this framework, we applied it on a recently proposed signature based hardware Trojan detection technique. The cross validation comparison of these learner shows that eager learners are able to detect the intrusion with 96% accuracy and also require less amount of memory and processing power compared to other machine learning techniques. Faiq Khalid, Imran Hafeez Abbassi, Osman Hasan, Falah R. Awwad, Syed Rafay Hasan |
ISCAS | 6 |
| 2016 | Analyzing Vulnerability of Asynchronous Pipeline to Soft Errors: Leveraging Formal Verification
Faiq Khalid, Syed Rafay Hasan, Osman Hasan, Falah R. Awwad |
J. Electron. Test. | 2 |
| 2016 | Clock domain crossing (CDC) in 3D-SICs: Semi QDI asynchronous vs loosely synchronous
Syed Rafay Hasan, Waqas Gul, Osman Hasan |
Integr. | 1 |
| 2014 | Abstracting Single Event Transient characteristics variations due to input patterns and fan-outabstractDue to shrinking feature sizes and significant reduction in noise margins, as CMOS technologies evolve toward ultra-deep sub-micron, digital circuits have become more susceptible to soft errors. Therefore, researchers have recently reported several approaches to model Single Event Transient (SET) propagation at gate or higher abstraction levels. However, contemporary techniques model only the possibility that SET pulse may be masked electrically, logically, or by time windowing. In this paper, the propagation induced pulse broadening (PIPB) phenomenon is further investigated and a new model which abstracts this phenomenon is proposed. This paper also investigates and abstracts the impact of input patterns and propagation paths on SET pulse width. Through electrical simulations, we validated our analysis. Ghaith Bany Hamad, Syed Rafay Hasan, Otmane Aït Mohamed, Yvon Savaria |
ISCAS | 2 |
| 2013 | A Library-Based Early Soft Error Sensitivity Analysis Technique for SRAM-Based FPGA Design
Claude Thibeault, Yassine Hariri, Syed Rafay Hasan, Christelle Hobeika, Yvon Savaria, Yves Audet, Fatima Zahra Tazi |
J. Electron. Test. | 3 |
| 2012 | A novel hybrid FIFO asynchronous clock domain crossing interfacing methodabstractMulti-clock domain circuits with Clock Domain Crossing (CDC) interfaces are emerging as an alternative to circuits with a global clock. CDC interfaces are susceptible to metastability, hence their design is very challenging. This paper presents a hybrid FIFO-asynchronous method for constructing robust CDC interfaces. The proposed design can handle arbitrary clock frequency ratios between the sender and receiver with random phase shifts. The proposed design avoids latency due to synchronizers with the asynchronous protocol modifications. Circuit simulation results confirm the operation and robustness of the design at maximum workloads, and arbitrary frequency ratios, over a temperature range of -50 to 50 degrees Celsius. The interface offers a maximum throughput of 606 million transfers per second without pausing the clock. Zaid Al-bayati, Otmane Aït Mohamed, Syed Rafay Hasan, Yvon Savaria |
ACM Great Lakes Symposium on VLSI | 3 |
| 2012 | Identification of soft error glitch-propagation paths: Leveraging SAT solversabstractIncrease in vulnerability to soft errors has affected the reliability of both synchronous and asynchronous circuits implemented in modern deep sub-micron technologies. Hence in such circuits, there is a growing need to identify the soft error glitch propagation possibility at an early stage in the design flow. This paper proposes a new methodology to obtain soft error glitch propagation paths in digital designs (both synchronous and asynchronous). To compute these paths, Multiway Decision Graphs (MDGs) and glitch-propagation sets (GP sets) are utilized in conjunction with Boolean Satisfiability solvers (MiniSat). The applicability of the proposed method is illustrated by implementing ISCAS89 benchmark sequential circuits, 8-bit adders, multipliers, and the Self-timed multiple-group pipeline asynchronous handshake circuits. The proposed SAT based methodology is on average 13 times faster than the best contemporary state-of-the-art techniques exhaustively analyze possible soft error glitch-propagation paths. Ghaith Bany Hamad, Otmane Aït Mohamed, Syed Rafay Hasan, Yvon Savaria |
ISCAS | 3 |
| 2011 | All digital skew tolerant synchronous interfacing methods for high-performance point-to-point communications in deep sub-micron SoCs
Syed Rafay Hasan, Normand Bélanger, Yvon Savaria, M. Omair Ahmad |
Integr. | 1 |
| 2009 | An All-digital Skew-adaptive Clock Scheduling Algorithm for Heterogeneous Multiprocessor Systems on Chips (MPSoCs)abstractIn this work, we propose a clock scheduling algorithm that is used to mitigate the effects of clock skew that can arise from thermal run-time variations. Depending on the amount of skew, the algorithm selects a different minimum delay tolerance value in order to correct the skew problems, without the performance penalties that are associated with static worst-case scheduling policies. The design was first implemented in MATLAB to obtain the data needed for the clock scheduling. Then, it was implemented in VHDL and synthesized using Xilinx's Virtex-II Pro technology library. Back annotated simulations prove the functionality of the proposed design. The adaptive scheduling scheme achieves up to 60% latency reduction, in our implemented example, compared to a static scheduling scheme. Syed Rafay Hasan, Bill Pontikakis, Yvon Savaria |
ISCAS | 1 |
| 2007 | Crosstalk Effects in Event-Driven Self-Timed Circuits Designed With 90nm CMOS TechnologyabstractSystems-on-chip (SoCs) designed in ultra-deep sub-micron technologies (90nm and beyond) often comprise modules in multiple clock domains (MCD), which are usually interconnected using asynchronous interfaces. At the same time, in ultra-deep sub-micron (DSM) technologies, minimum width, spacing, inter-metal dielectric lengths are reduced, as well as distances between metal layers. These trends raise the coupling capacitance resulting in more severe crosstalks. Therefore, asynchronous interfaces may be subject to crosstalk in ultra-DSM technologies. In this paper, a quantitative investigation is performed to approximate the crosstalk effects in 90nm technology, and to compare them with effects in other DSM technologies. It is found that for wire lengths of 1mm, and more, crosstalk effects in a 90nm technology are substantially higher, about 1.3 times, than in a 180nm technology. Furthermore, three well known self-timed asynchronous design methods are analyzed with regards to crosstalk and the importance of coupling capacitances is established. It is shown that glitches can cause errors in self-timed designs. To our knowledge, this paper is the first to report crosstalk sensitivity in self-timed circuits, which are notably proposed as a solution to the timing problems found in advanced SoCs. Syed Rafay Hasan, Yvon Savaria |
ISCAS | 1 |
| 2004 | Optimal partitioning of globally asychronous locally synchronous processor arraysabstractWith ongoing advances of semiconductor technology, power dissipation has been moving higher on the list of VLSI design constraints. In most high-performance synchronous VLSI designs, the distribution of low-skew global clock signals approaching GigaHertz range is the single largest source of power consumption. GALS design style offers a solution to this issue by dividing synchronous design into smaller locally synchronous sub-blocks. Smaller sub-blocks reduce capacitance in clock distribution networks because they need less H-tree levels. However, this implies a large number of sub-blocks, which increases the asynchronous power overhead. This work investigates these GALS power tradeoffs. This is, to our knowledge, the first paper to propose closed form models for optimum number of partitions that gives minimum power for a GALS array of identical processors. The models can serve as a useful firsthand guideline for designers in initial design stages. Experimental results verify the effectiveness of the model. Adhir Upadhyay, Syed Rafay Hasan, Mohamed Nekili |
ACM Great Lakes Symposium on VLSI | 2 |