Shahzad Muzaffar

dblp:98/4719 · DBLP profile ↗
← Back
14ranked-venue papers
11as first author
3since 2021 · last 2024
0000-0001-9237-1148ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 10 first-author · 3 since 2021Computer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2024 OSHDA: A Containerized CAD Tool for the Design and Analysis of Behavioral FSM Logic Locking
abstract
This paper introduces the Open-source Secure Hardware Design and Analysis (OSHDA) toolchain for the logic locking of finite-state machines (FSMs) at the behavioral level. OSHDA's FSM obfuscation method is based on the recently developed State Permutation Logic Locking (SPeLL) algorithm which obfuscates the behavioral transition graph of the FSM, thus avoiding the use of dummy states and reducing exposure to reverse engineering attacks. In addition to implementing the SPeLL algorithm, the toolchain implements a full logic synthesis flow, including the evaluation of the gate-level SPeLL hardware overhead for both FPGA and ASIC designs. In particular, OSHDA enables the automation of trade-off analysis between the strength of SPeLL security and its hardware overhead. The paper further describes attempted attacks on SPeLL using state-of-the-art de-obfuscation tools and identifies research gaps in behavioral de-obfuscation that must be addressed before one can successfully de-obfuscate SPeLL. OSHDA comes with its own scripting subsystem for augmenting its analysis, adding de-obfuscation methods, and integrating physical design tools. Finally, OSHDA is deployed as a hardware security microservice using the Docker framework.
Esrat Khan, Shahzad Muzaffar, Lamees M. Al Qassem, Ibrahim M. Elfadel
VLSI-SoC2
2022 Logic Locking of Finite-State Machines Using Transition Obfuscation
abstract
In this paper, we introduce a novel algorithm for securing sequential circuits at the Register-Transfer Level (RTL) that does not require any state augmentation. The algorithm is based on the encryption of the state encodings with a key that is known only to the IP provider. When the correct key is input at runtime, the sequential circuit will operate as designed, otherwise it will operate according to a state transition map that is defined by the wrong key. We call this mode of operation: transition obfuscation. One important advantage of the proposed method is that using the wrong key does not necessarily result in the sequential circuit getting stuck at any one state or getting trapped within any black hole. As a result, the secured sequential circuit is more immune to reverse engineering attacks, and because of the large number of wrong full-state transition maps, more immune to side-channel attacks. A full low-complexity, RTL design methodology based on the new algorithm is presented along with extensive experiments quantifying its design overhead and illustrating its advantages in terms of immunity to reverse engineering attacks.
Shahzad Muzaffar, Ibrahim M. Elfadel
VLSI-SoC1
2021 Beyond Arduino: A Guide for the Perplexed
abstract
Arduino IDE and boards have been used in training, hobby projects, and lab experiments because of their ease of use, flexibility, and readily available libraries. However, they cannot be used for commercial products and customized systems that target low-cost and low-power operation, are resource-constrained and need to comply with industrial and regulatory standards. This paper provides a methodology and guidelines on how to migrate a lab prototype developed using the Arduino environment and to a near-product prototype using an industry- strength IDE and MDK development environment. The guidelines involve a sequence of steps whose main goal is to minimize the hardware and software debugging effort. Each step is focused on bringing one aspect under active development while keeping everything else fixed and bug-free. Moreover, at each step, a working reference is set up to enable cross-checks in case of any unexpected problem. The guidelines are illustrated with a case study from our own work in which an Arduino prototype has been successfully transformed into a wearable with an extremely small footprint. In addition to their value for embedded system design, these guidelines have the educational value of contrasting the learning outcomes of embedded system courses based on the Arduino framework vs. those based on an industry-driven MDK and IDE.
Shahzad Muzaffar, Ibrahim M. Elfadel
ISCAS1
2020 Lessons Learned the Hard Way
abstract
“Fail often to succeed sooner” is a common mantra that we are told is the secret to success. When reporting research results, however, scholars rarely write about their failed attempts and only focus on the successful ones. Perhaps the source of this disconnect between what we preach and what we do can be found in the underlying assumption that published work is meant to move the field forward and failed attempts supposedly do not. The goal of the confessions presented in this paper is to show that even failed attempts are genuine and valuable contributions to our field provided that we learn from our mistakes and correct them. The 27 confessions span from planning oversights, digital and analog design errors, misunderstanding of devices, overlooked parasitics, LVS errors, and troubles in testing.
Tobi Delbruck, Ibrahim M. Elfadel, Shahzad Muzaffar, Germain Haessig, Bo Wang 0012, Amine Bermak, Rui Graca, Luis A. Camuñas-Mesa, Bathiya Senevirathna, Pamela Abshire, Bernabé Linares-Barranco, Saeed Afshar, Shih-Chii Liu, Runchun Wang, Piotr Dudek, Stephen J. Carey, José M. de la Rosa 0001, Marc Dandin, Sheung Lu, Vincent Frick, Teresa Serrano-Gotarredona, Paula López Martinez 0001, Melika Payvand, Advait Madhavan, Eric R. Fossum, Juan Camilo Vasquez Tieck, Yan Liu 0016, Timothy G. Constandinou, Alexander Serb, Ricardo Carmona-Galán, Robert Nawrocki, Walter D. Leon-Salas
ISCAS3
2020 An Inference Hardware Accelerator for EEG-Based Emotion Detection
abstract
The wearability of emotion classifiers is a must if they are to significantly improve the social integration of patients suffering from neurological disorders. Such wearability requires the use of low-power hardware accelerators that would enable near real-time classification and extended periods of operations. In this paper, we architect, design, implement, and test a handcrafted, hardware Convolutional Neural Network, named BioCNN, optimized for EEG-based emotion detection and other similar bio-medical applications. The architecture of BioCNN is based on aggressive pipelining and hardware parallelism that maximizes resource re-use and minimizes memory footprint. The FEXD and DEAP datasets are used to test the BioCNN prototype that is implemented using the Digilent Atlys Board with a low-cost Spartan-6 FPGA. The experimental results show that BioCNN has a competitive energy efficiency of 11GOps/W, a throughput of 1.65GOps that is in line with the real-time specification of a wearable device, and a latency of less than 1ms, which is much smaller than the 150ms required for human interaction times. Its emotion inference accuracy is competitive with the top software-based emotion detectors.
Hector A. Gonzalez, Shahzad Muzaffar, Jerald Yoo, Ibrahim M. Elfadel
ISCAS2
2020 Dynamic Edge-coded Protocols for Low-power, Device-to-device Communication
abstract
Clock and Data Recovery (CDR) has been a foundational receiver component in serial communications. Yet this component is known to add significant design complexity to the receiver and to consume significant resources in area and power. In the resource-limited world of constrained IoT nodes, the need of including CDR in the communication link is being re-assessed and new techniques for achieving reliable serial transmission without CDR have been emerging. These new techniques are distinguished by their use of transition edges rather than bit times for coding and detection. This article presents the design, implementation, and testing of a novel CDR-less transmission protocol that achieves significant improvements in data rate, reliability, packet security, and power efficiency with respect to state-of-the-art CDR-less techniques. The new protocol further tolerates significant jitters and clock discrepancies between transmitter and receiver. An FPGA and an ASIC (65 nm technology) implementation of the protocol have shown it to consume around 19μ W of power at a clock rate of 25 MHz, and to have a small footprint with a gate count of approximately 2,098 gates. In particular, the new protocol reduces area by more than 87% and power by more than 78% in comparison with CDR-based serial bit transfer protocols. Furthermore, the new protocol is shown to be versatile in its applications to available communication media, including wired, wireless, infrared, and human-body channels, under a variety of digital modulation schemes.
Shahzad Muzaffar, Ibrahim M. Elfadel
ACM Trans. Sens. Networks1
2019 Double Data Rate Dynamic Edge-Coded Signaling for Low-Power IoT Communication
abstract
Dynamic Edge-Coded Signaling (ECS) is a recently introduced protocol for single-channel signaling between constrained IoT nodes. One of the most distinguishing features of ECS is that its receiver does not require any circuitry for clockand-data recovery. Other important ECS features include its tolerance with respect to clock variations between transmitter and receiver and its amenability to seamlessly integrate lightweight cryptographic algorithms to ensure secure communication. ECS encodes information using pulse counts with the counting based on one of the pulse edges. In this paper, we address the problem of improving the ECS data rate for a given clock frequency and under a given power envelop by using both pulse edges of the ECS pulse stream. We call the novel protocol double-data-rate ECS (DDR-ECS) in analogy with DDR memory systems. While the concept is intuitive and attractive, its hardware implementation is not. This paper, therefore, presents an efficient hardware design of the DDR-ECS transceiver that preserves the ECS built-in features while essentially doubling the data rate at the same clock frequency and within the same power budget. A 65nm ASIC synthesis of the transceiver shows that DDR-ECS consumes the ECS equivalent power of 19μW, uses a small form factor of only 1934 gates, and doubles the dynamic data rate to the 7.8-44.4 Mb/s range with an average of 12 Mb/s at a clock rate of 25MHz.
Shahzad Muzaffar, Ibrahim M. Elfadel
VLSI-SoC1
2019 A Domain-Specific Processor Microarchitecture for Energy-Efficient, Dynamic IoT Communication
abstract
In this paper, we present a domain-specific processor architecture, named pulsed-index communication interface architecture (PICIA), for single-channel IoT communication based on the recently introduced pulsed-signaling protocols, according to which information is encoded as series of pulses representing ON bits. In addition to the traditional aspects of instruction set architecture (ISA) design such as addressing modes, instruction types, instruction formats, registers, interrupts, and external I/O, the ISA includes domain-specific instructions that facilitate bit stream encoding and decoding based on the pulsed-signaling techniques. The domain-specific PICIA microarchitecture employs a set of optimized processing blocks that can be used programmatically to encode and decode the transmitted data in the most economical way. The PICIA allows customizations that support both standard pulsed-signaling techniques and specialized protocols that belong to the same family. The PICIA design further allows an amalgamation of software and hardware that significantly reduces the number of instructions required to implement a given communication interface without impacting the data rates and reliability of the pulsed-signaling protocols. The PICIA processor has been implemented in Verilog HDL and tested using a Xilinx Spartan-6 field-programmable gate array (FPGA). Furthermore, a 65-nm application-specific integrated circuits (ASIC) synthesis of the design confirms the small-footprint and low-power features of PICIA. The consumed power has been evaluated at 31.14μW with an energy efficiency of less than 10 pJ/bit.
Shahzad Muzaffar, Ibrahim M. Elfadel
IEEE Trans. Very Large Scale Integr. Syst.1
2018 An Instruction Set Architecture for Low-power, Dynamic IoT Communication
abstract
This paper presents an instruction set architecture (ISA) dedicated to the rapid and efficient implementation of single-channel IoT communication interfaces. The architecture is meant to provide a programming interface for the implementation of signaling protocols based on the recently introduced pulsed-index schemes. In addition to the traditional aspects of ISA design such as addressing modes, instruction types, instruction formats, registers, interrupts, and external I/O, the ISA includes special-purpose instructions that facilitate bit stream encoding and decoding based on the pulsed-index techniques. Verilog HDL is used to synthesize a fully functional processor based on this ISA and provide both an FPGA implementation and a synthesised ASIC design in GLOBALFOUNDRIES 65nm. The ASIC design confirms the low-power features of this ISA with consumed power around 31µW and energy efficiency of less than 10pJ/bit.
Shahzad Muzaffar, Ibrahim M. Elfadel
VLSI-SoC1
2017 A pulsed decimal technique for single-channel, dynamic signaling for IoT applications
abstract
Pulsed-Index Communication (PIC) is a recent technique for single-channel communication which is based on the principle of transferring the indices of only the ON bits in the form of a series of pulse streams. In this paper, we present a modified version of PIC which is based on the same underlying idea but with key improvements in data rate and reliability. The proposed technique is called Pulsed Decimal Communication (PDC). Like PIC, PDC is a protocol for single-channel, high-data rate, low-power dynamic signaling that does not require any clock and data recovery. It however achieves higher data rates by introducing a three-step algorithm, comprising a segmentation, an encoding, and a sub-segmentation step. The segmentation step is used to split the data word into smaller segments and therefore smaller decimal numbers to represent them. The encoding step reduces the number of ON bits in the data and relocates them to lower indices. The sub-segmentation step is used to split further the segments into smaller sub-segments. The complete process reduces the number of pulses required to transmit binary data, thus improving the data rate. Compared with PIC, PDC achieves a 78% improvement in data rate and is more reliable as it eliminates the variations in the number of symbols to be transmitted. An FPGA and an ASIC (65nm technology) implementation of the protocol show that the low-power operation and small footprint of PIC are maintained in PDC, which consumes around 25of power at a clock frequency of 25MHz with a gate count of approximately 2150.
Shahzad Muzaffar, Ibrahim M. Elfadel
VLSI-SoC1
2016 Automatic protocol configuration in single-channel low-power dynamic signaling for IoT devices
abstract
Pulsed-Index Communication (PIC) is a novel technique for single-channel, high-data rate, low-power dynamic signaling that does not require any clock and data recovery. It is fully adapted to the simple yet robust communication needs of Internet of Things (IoT) devices and sensors. However, its error-free operation with maximum data rate requires a careful and judicious setting of PIC data packet and pulse timing parameters. In this paper, we present a new algorithm for automatically detecting and setting the PIC protocol parameters at the power-on phase while removing the restriction on the IoT devices in the PIC network to communicate at a baud rate. The hardware realization of the algorithm is power-efficient and uses closed-form formulas that assign suitable protocol parameters to both ends of the transmission link based on clock rate differences. This difference is determined by a preliminary exchange of clock pulse streams between the transmitter and the receiver. The automatic parameter setting remains operational even in the presence of variations between the local clock frequencies of the IoT devices communicating via PIC. The algorithm is illustrated in the case of several IoT devices with different local clock frequencies that are in need to synchronize their communication parameters with respect to the clock frequency of a master gateway node. A power-on PIC parameter configuration process is rigorously specified, and both an FPGA and an ASIC implementations are presented. In particular, we show that for an ASIC implementation in 65nm technology, the low-power operation of PIC is maintained, consuming only 4.35µW of power at a clock frequency of 25MHz. This architecture is experimentally verified and tested on a point-to-point communication link between two IoT devices connected via a single PIC channel in a master-slave mode.
Shahzad Muzaffar, Numan Saeed, Ibrahim M. Elfadel
VLSI-SoC1
2015 A pulsed-index technique for single-channel, low-power, dynamic signaling
Shahzad Muzaffar, Jerald Yoo, Ayman Shabra, Ibrahim M. Elfadel
DATE1
2015 Power management of pulsed-index communication protocols
abstract
Pulsed-Index Communication (PIC) is a novel technique for single-channel, high-data-rate, low-power dynamic signaling that does not require any clock and data recovery (CDR). It is fully adapted to the simple yet robust communication needs of IoT devices and sensors. Prior work has focused on the power savings that this protocol can achieve as a result of the elimination of circuitry devoted to clock and data recovery. In this paper, we show that further power saving can be achieved using the duty cycle of the pulse as a power control parameter. This power control policy is applied to a single-wire link with significant power saving achieved above and beyond the savings due to the CDR elimination. These power savings are obtained without any impact on data rate. The pulse control policy is implemented using 45nm CMOS technology and verified on various, single-channel communication links.
Shahzad Muzaffar, Ibrahim M. Elfadel
ICCD1
2015 Timing and robustness analysis of Pulsed-Index protocols for single-channel IoT communications
abstract
Pulsed-Index Communication (PIC) is a novel technique for single-channel, high-data-rate, low-power dynamic signaling that does not require any clock and data recovery. It is fully adapted to the simple yet robust communication needs of IoT devices and sensors. In this paper, we present a full quantitative analysis of the timing and robustness properties of PIC protocols, including the impact of important protocol parameters such as pulse width and inter-symbol delays on average data rate and protocol robustness with respect to clock variations. The main result of this paper is a theoretical upper bound on clock variability between transmitter and receiver below which the protocol operates with zero decoding error over an ideal channel. This bound is verified experimentally using a full FPGA implementation that includes point-to-point transmission between two TI MSP430 microcontrollers, acting as two IoT sensor nodes over a single-wire connection.
Shahzad Muzaffar, Ibrahim M. Elfadel
VLSI-SoC1