Rehan Hafiz

dblp:10/9281 · DBLP profile ↗
← Back
19ranked-venue papers
1as first author
2since 2021 · last 2022
0000-0002-5062-3068ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13Artificial intelligence and machine learning · 3 · 2 since 2021Software engineering, systems software and programming languages · 3Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
8 papers
Emerging computing paradigms · 46% Integrated circuit design · 17% Hardware accelerators and domain-specific architectures · 13%

Topics — the 17 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Emerging computing paradigms
approximate computing
2.072020
PEMACx: A Probabilistic Error Analysis Methodology for Adders with Cascaded Approximate Units · DAC 2020
Probabilistic Error Modeling for Approximate Adders · IEEE Trans. Computers 2017
Probabilistic Error Analysis of Approximate Recursive Multipliers · IEEE Trans. Computers 2017
Emerging computing paradigms › approximate computing › approximate circuit design › approximate arithmetic circuits
approximate adder
1.352020
PEMACx: A Probabilistic Error Analysis Methodology for Adders with Cascaded Approximate Units · DAC 2020
Probabilistic Error Modeling for Approximate Adders · IEEE Trans. Computers 2017
QuAd: Design and Analysis of Quality-Area Optimal Low-Latency Approximate Adders · DAC 2017
Memory systems › data locality
data reuse
0.412020
SuperSlash: A Unified Design Space Exploration and Model Compression Methodology for Design of Deep Learning Accelerators With Reduced Off-Chip Memory Access Volume · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Electronic design automation
design space exploration
0.412020
SuperSlash: A Unified Design Space Exploration and Model Compression Methodology for Design of Deep Learning Accelerators With Reduced Off-Chip Memory Access Volume · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.412020
SuperSlash: A Unified Design Space Exploration and Model Compression Methodology for Design of Deep Learning Accelerators With Reduced Off-Chip Memory Access Volume · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Hardware accelerators and domain-specific architectures
model compression
0.412020
SuperSlash: A Unified Design Space Exploration and Model Compression Methodology for Design of Deep Learning Accelerators With Reduced Off-Chip Memory Access Volume · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Memory systems › memory access
off-chip memory access
0.412020
SuperSlash: A Unified Design Space Exploration and Model Compression Methodology for Design of Deep Learning Accelerators With Reduced Off-Chip Memory Access Volume · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Integrated circuit design › digital circuit design
arithmetic circuit design
0.432017
A low latency generic accuracy configurable adder · DAC 2015
Probabilistic Error Modeling for Approximate Adders · IEEE Trans. Computers 2017
Probabilistic Error Analysis of Approximate Recursive Multipliers · IEEE Trans. Computers 2017
Integrated circuit design
low-power circuit design
0.422017
QuAd: Design and Analysis of Quality-Area Optimal Low-Latency Approximate Adders · DAC 2017
Invited - Cross-layer approximate computing: from logic to architectures · DAC 2016
Integrated circuit design › digital circuit design › arithmetic circuit design
adder design
0.322017
A low latency generic accuracy configurable adder · DAC 2015
Probabilistic Error Modeling for Approximate Adders · IEEE Trans. Computers 2017
Integrated circuit design
digital circuit design
0.322016
A low latency generic accuracy configurable adder · DAC 2015
An area-efficient consolidated configurable error correction for approximate hardware accelerators · DAC 2016
Emerging computing paradigms › approximate computing
approximate multiplier
0.312017
Probabilistic Error Analysis of Approximate Recursive Multipliers · IEEE Trans. Computers 2017
Emerging computing paradigms › approximate computing
cross-layer approximate computing
0.212016
Invited - Cross-layer approximate computing: from logic to architectures · DAC 2016
Hardware reliability and fault tolerance
error correction
0.212016
An area-efficient consolidated configurable error correction for approximate hardware accelerators · DAC 2016
Integrated circuit design › digital circuit design › arithmetic circuit design
multiplier design
0.112017
Probabilistic Error Analysis of Approximate Recursive Multipliers · IEEE Trans. Computers 2017
Emerging computing paradigms › approximate computing › approximate circuit design
approximate arithmetic circuits
0.112016
Invited - Cross-layer approximate computing: from logic to architectures · DAC 2016
Hardware reliability and fault tolerance
soft errors
0.112015
A low latency generic accuracy configurable adder · DAC 2015

Methods — techniques the papers use, named apart from their topics

probability mass function · 0.6recursive carry-out probability · 0.4ranking function · 0.4pruning · 0.4probability mass function computation · 0.4layer fusion · 0.4probabilistic modeling · 0.3mathematical analysis · 0.3closed-form analysis · 0.3RTL modeling · 0.3
YearPublicationVenuePosition
2022 FogAdapt: Self-supervised domain adaptation for semantic segmentation of foggy images
Javed Iqbal 0007, Rehan Hafiz, Mohsen Ali
Neurocomputing2
2022 Distribution regularized self-supervised learning for domain adaptation of semantic segmentation
Javed Iqbal 0007, Hamza Rawal, Rehan Hafiz, Yu-Tseh Chi, Mohsen Ali
Image Vis. Comput.3
2020 PEMACx: A Probabilistic Error Analysis Methodology for Adders with Cascaded Approximate Units
abstract
In this paper, we propose a novel methodology for efficiently computing the Probability Mass Function (PMF) of error at the output of a major class of approximate adders that comprise of a cascade of multiple stages of smaller approximate adder units. The proposed methodology utilizes the carry-out probability of the previous stage along with the input probabilities of the current stage to recursively computes the PMF of error of the current stage. The proposed methodology eliminates the need for exhaustive simulations and, therefore, can be used for efficiently analyzing error distribution of a wide variety of low-power large bit-width adders with cascaded approximate adder units. Experimental results demonstrate that the proposed methodology provides error estimates that commensurate with exhaustive simulations, while offering a speedup of at least 2958x for 8-bit (or larger) approximate adder configurations.
Muhammad Abdullah Hanif, Rehan Hafiz, Osman Hasan, Muhammad Shafique 0001
DAC2
2020 SuperSlash: A Unified Design Space Exploration and Model Compression Methodology for Design of Deep Learning Accelerators With Reduced Off-Chip Memory Access Volume
abstract
Deploying deep learning (DL) models on resource-constrained embedded devices is a challenging task. The limited on-chip memory on such devices results in increased off-chip memory access volume, thus limiting the size of DL models that can be efficiently realized in such systems. Design space exploration (DSE) under memory constraint, or to achieve minimal off-chip memory access volume, has recently received much attention. Unfortunately, DSE alone cannot reduce the amount of off-chip memory accesses beyond a certain point due to the fixed model size. Model compression via pruning can be employed to reduce the size of the model and the associated off-chip memory accesses. However, in this article, we demonstrate that pruned models with even the same accuracy and model size may require a different number of off-chip memory accesses depending upon the pruning strategy adopted. Thus, mainstream pruning techniques may not be closely tied to the design goals, and thereby hard to be integrated with existing DSE techniques. To overcome this problem, we propose SuperSlash, a unified solution for DSE and model compression. SuperSlash estimates off-chip memory access volume overhead of each layer of a DL model by exploring multiple design candidates. In particular, it evaluates multiple data reuse strategies for each layer, along with the possibility of layer fusion. Layer fusion aims at reducing the off-chip memory access volume by avoiding the intermediate off-chip storage of a layer's output and directly using it for processing of the subsequent layer. SuperSlash then guides the pruning process via a ranking function, which ranks each layer according to its explored off-chip memory access cost. We demonstrate that SuperSlash not only offers an extensive design space coverage but also provides lower off-chip memory access volume (up to 57.71%, 25.83%, 47.73%, and 29.02% reduction for VGG16, ResNet56, ResNet110, and MobileNetV1, respectively) as compared to the state-of-art.
Hazoor Ahmad, Tabasher Arif, Muhammad Abdullah Hanif, Rehan Hafiz, Muhammad Shafique 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2018 An overview of next-generation architectures for machine learning: Roadmap, opportunities and challenges in the IoT era
abstract
The number of connected Internet of Things (IoT) devices are expected to reach over 20 billion by 2020. These range from basic sensor nodes that log and report the data to the ones that are capable of processing the incoming information and taking an action accordingly. Machine learning, and in particular deep learning, is the de facto processing paradigm for intelligently processing these immense volumes of data. However, the resource inhibited environment of IoT devices, owing to their limited energy budget and low compute capabilities, render them a challenging platform for deployment of desired data analytics. This paper provides an overview of the current and emerging trends in designing highly efficient, reliable, secure and scalable machine learning architectures for such devices. The paper highlights the focal challenges and obstacles being faced by the community in achieving its desired goals. The paper further presents a roadmap that can help in addressing the highlighted challenges and thereby designing scalable, high-performance, and energy efficient architectures for performing machine learning on the edge.
Muhammad Shafique 0001, Theocharis Theocharides, Christos-Savvas Bouganis, Muhammad Abdullah Hanif, Faiq Khalid, Rehan Hafiz, Semeen Rehman
DATE6
2018 Error resilience analysis for systematically employing approximate computing in convolutional neural networks
abstract
Approximate computing is an emerging paradigm for error resilient applications as it leverages accuracy loss for improving power, energy, area, and/or performance of an application. The spectrum of error resilient applications includes the domains of Image and video processing, Artificial Intelligence (AI) and Machine Learning (ML), data analytics, and other Recognition, Mining, and Synthesis (RMS) applications. In this work, we address one of the most challenging question, i.e., how to systematically employ approximate computing in Convolution Neural Networks (CNNs), which are one of the most compute-intensive and the pivotal part of AI. Towards this, we propose a methodology to systematically analyze error resilience of deep CNNs and identify parameters that can be exploited for improving performance/efficiency of these networks for inference purposes. We also present a case study for significance-driven classification of filters for different convolutional layers, and propose to prune those having the least significance, and thereby enabling accuracy vs. efficiency tradeoffs by exploiting their resilience characteristics in a systematic way.
Muhammad Abdullah Hanif, Rehan Hafiz, Muhammad Shafique 0001
DATE2
2017 QuAd: Design and Analysis of Quality-Area Optimal Low-Latency Approximate Adders
abstract
Approximate circuits exploit error resilience property of applications to tradeoff computation quality (accuracy) for gaining advantage in terms of performance, power, and/or area. While state-of-the-art low-latency approximate adders provide an accuracy-area-latency configurable design space, the selection of a particular configuration from the design space is still manually done. In this paper, we analytically analyze different structural properties of low-latency approximate adders to formulate a new adder model, Quality-area optimal Low-Latency approximate Adder (QuAd). It provides an increased design space as compared to state-of-the-art, providing design points that require less logic area for the same accuracy, as compared to state-of-the-art approximate adders. Furthermore, based upon our mathematical analysis, we show that, provided a latency constraint, an adder configuration with the highest quality and lowest area requirement can effortlessly be selected from the whole design space of QuAd adder model, without requiring any optimization strategy or numerical simulation. Our experimental results validate the developed model and also the quality-area optimality of our optimal QuAd adder configuration. For functional verification and prototyping, we have used a Xilinx Virtex-6 FPGA. RTL/behavioral models and MATLAB equivalent scripts, of our proposed adder model are made open source, to facilitate further research and development.
Muhammad Abdullah Hanif, Rehan Hafiz, Osman Hasan, Muhammad Shafique 0001
DAC2
2017 Embracing approximate computing for energy-efficient motion estimation in high efficiency video coding
abstract
Approximate Computing is an emerging paradigm for developing highly energy-efficient computing systems. It leverages the inherent resilience of applications to trade output quality with energy efficiency. In this paper, we present a novel approximate architecture for energy-efficient motion estimation (ME) in high efficiency video coding (HEVC). We synthesized our designs for both ASIC and FPGA design flows. ModelSim gate-level simulations are used for functional and timing verification. We comprehensively analyze the impact of heterogeneous approximation modes on the power/energy-quality tradeoffs for various video sequences. To facilitate reproducible results for comparisons and further research and development, the RTL and behavioral models of approximate SAD architectures and constituting approximate modules are made available at https://sourceforge.net/projects/lpaclib/.
Walaa El-Harouni, Semeen Rehman, Bharath Srinivas Prabakaran, Akash Kumar 0001, Rehan Hafiz, Muhammad Shafique 0001
DATE5
2017 Confined projection on selected sub-surface using a robust binary-coded pattern for pico-projectors
Shafaq Mussadiq, Rehan Hafiz, Muhammad Abdullah Jamal
Multim. Syst.2
2017 Probabilistic Error Analysis of Approximate Recursive Multipliers
abstract
Approximate multipliers are gaining importance in energy-efficient computing and require careful error analysis. In this paper, we present the error probability analysis for recursive approximate multipliers with approximate partial products. Since these multipliers are constructed from smaller approximate multiplier building blocks, we propose to derive the error probability in an arbitrary bit-width multiplier from the probabilistic model of the basic building block and the probability distributions of inputs. The analysis is based on common features of recursive multipliers identified by carefully studying the behavioral model of state-of-the-art designs. By building further upon the analysis, Probability Mass Function (PMF) of error is computed by individually considering all possible error cases and their inter-dependencies. We further discuss the generalizations for approximate adder trees, signed multipliers, squarers and constant multipliers. The proposed analysis is validated by applying it to several state-of-the-art approximate multipliers and comparing with corresponding simulation results. The results show that the proposed analysis serves as an effective tool for predicting, evaluating and comparing the accuracy of various multipliers. Results show that for the majority of the recursive multipliers, we get accurate error performance evaluation. We also predict the multipliers' performance in an image processing application to demonstrate its practical significance.
Sana Mazahir, Osman Hasan, Rehan Hafiz, Muhammad Shafique 0001
IEEE Trans. Computers3
2017 Probabilistic Error Modeling for Approximate Adders
abstract
Approximate adders are widely being advocated as a means to achieve performance gain in error resilient applications. In this paper, a generic methodology for analytical modeling of probability of occurrence of error and the Probability Mass Function (PMF) of error value in a selected class of approximate adders is presented, which can serve as performance metrics for the comparative analysis of various adders and their configurations. The proposed model is applicable to approximate adders that comprise of subadder units of uniform as well as non-uniform lengths. Using a systematic methodology, we derive closed form expressions for the probability of error for a number of state-of-the-art high-performance approximate adders. The probabilistic analysis is carried out for arbitrary input distributions. It can be used to study the dependence of error statistics in an adder's output on its configuration and input distribution. Moreover, it is shown that by building upon the proposed error model, we can estimate the probability of error in circuits with multiple approximate adders. We also demonstrate that, using the proposed analysis, the comparative performance of different approximate adders can be correctly predicted in practical applications of image processing.
Sana Mazahir, Osman Hasan, Rehan Hafiz, Muhammad Shafique 0001, Jörg Henkel
IEEE Trans. Computers3
2016 An area-efficient consolidated configurable error correction for approximate hardware accelerators
abstract
Approximate adders are widely being advocated for developing hardware accelerators to perform complex arithmetic operations. Most of the state-of-the-art accuracy configurable approximate adders utilize some integrated Error Detection and Correction (EDC) circuitry. Consequently, the accumulated area overhead due to the EDC (integrated within individual adders) is significant. In this paper, we propose a low-cost Consolidated Error Correction (CEC) unit, that essentially corrects the accumulated error at the accelerator output. The proposed CEC is based on a mathematical model of approximation error. We integrate our CEC unit in approximate hardware accelerators deployed in different applications to demonstrate its area savings and speed enhancement compared to state-of-the-art.
Sana Mazahir, Osman Hasan, Rehan Hafiz, Muhammad Shafique 0001, Jörg Henkel
DAC3
2016 Invited - Cross-layer approximate computing: from logic to architectures
abstract
We present a survey of approximate techniques and discuss concepts for building power-/energy-efficient computing components reaching from approximate accelerators to arithmetic blocks (like adders and multipliers). We provide a systematical understanding of how to generate and explore the design space of approximate components, which enables a wide-range of power/energy, performance, area and output quality tradeoffs, and a high degree of design flexibility to facilitate their design. To enable cross-layer approximate computing, bridging the gap between the logic layer (i.e. arithmetic blocks) and the architecture layer (and even considering the software layers) is crucial. Towards this end, this paper introduces open-source libraries of low-power and high-performance approximate components. The elementary approximate arithmetic blocks (adder and multiplier) are used to develop multi-bit approximate arithmetic blocks and accelerators. An analysis of data-driven resilience and error propagation is discussed. The approximate computing components are a first steps towards a systematic approach to introduce approximate computing paradigms at all levels of abstractions.
Muhammad Shafique 0001, Rehan Hafiz, Semeen Rehman, Walaa El-Harouni, Jörg Henkel
DAC2
2016 Automatic selection of color reference image for panoramic stitching
Muhammad Twaha Ibrahim, Rehan Hafiz, Muhammad Murtaza Khan, Yongju Cho
Multim. Syst.2
2015 A low latency generic accuracy configurable adder
abstract
High performance approximate adders typically comprise of multiple smaller sub-adders, carry prediction units and error correction units. In this paper, we present a low-latency generic accuracy configurable adder to support variable approximation modes. It provides a higher number of potential configurations compared to state-of-the-art, thus enabling a high degree of design flexibility and trade-off between performance and output quality. An error correction unit is integrated to provide accurate results for cases where high accuracy is required. Furthermore, an associated scheme for error probability estimation allows convenient comparison of different approximate adder configurations without requiring the need to numerically simulate the adder. Our experimental results validate the developed error model and also the lower latency of our generic accuracy configurable adder over state-of-the-art approximate adders. For functional verification and prototyping, we have used a Xilinx Virtex-6 FPGA. Our adder model and synthesizable RTL are made open-source.
Muhammad Shafique 0001, Rehan Hafiz, Jörg Henkel
DAC3
2015 Stabilization of panoramic videos from mobile multi-camera platforms
Rehan Hafiz, Muhammad Murtaza Khan, Yongju Cho, Jihun Cha
Image Vis. Comput.2
2014 Power & throughput optimized lifting architecture for Wavelet Packet Transform
abstract
This paper presents area-power efficient architectures for the lifting based Wavelet Packet Transform (WPT). Using Daubechies 6 as an example, three different approaches to the lifting scheme implementation are optimized. For higher level decompositions, a novel Fibonacci based technique to optimally compute the number of processing elements per level is presented. Comparisons between FPGA implementations of various architectures show a throughput-to-power ratio improvement of 62% over previously implemented WPT architectures. The architecture consumes a smaller area, while consuming a dynamic power of 46mW at a maximum throughput of 342 Mbits/sec per level.
Masab Ahmad, Awais M. Kamboh, Rehan Hafiz
ISCAS3
2013 ISOMER: integrated selection, partitioning, and placement methodology for reconfigurable architectures
abstract
Quality system design on dynamic partially reconfigurable platform needs exploration of a vast and multidimensional design space for (1) selection among implementation variants of hardware accelerators, (2) partitioning the reconfigurable fabric, and (3) their placement on the reconfigurable fabric partitions. This paper presents a novel methodology ISOMER for integrated solution of selection, partitioning and placement for performance optimization. Architecture under consideration is a general purpose processor coupled with reconfigurable fabric that can be partitioned in multi-sized partially reconfigurable bins. Our methodology determines performance-efficient partitioning and usage of reconfigurable fabric. Extensive evaluation illustrates that our methodology is scalable and outperforms state-of-the-art techniques for non-partially reconfigurable architectures.
Rana Muhammad Bilal, Rehan Hafiz, Muhammad Shafique 0001, Saad Shoaib, Asim Munawar, Jörg Henkel
ICCAD2
2004 An optimized hardware accelerator for real time registration of aerial video imagery and its applications
abstract
Aerial video imagery is widely used in mapping, surveillance and monitoring applications. An aerial video can give sufficient information, but it does not offer the freedom and flexibility of working with a geo-registered image or map. Our paper provides a cost effective, robust and efficient solution for real time video registration. We propose a fast compact FPGA processing module. The FPGA module implements the computationally intensive routines of the algorithm. Three level pipelining is incorporated to enhance the clock speed to 30 MHz. We also present two applications of our proposed hardware using the core registration algorithm. The first is an air-to-ground target tracker. The second application is a jerky hand-held video stabilizer application. It can be used for improved targeting from telescope-mounted weapons e.g. sniper, while firing from a moving vehicle. C++ implementations of these algorithms have been tested to demonstrate encouraging results.
Rehan Hafiz, M. Ali Tahir, Omair Arshad, Shoab Ahmad Khan
ICIG1