VLDB 2026 Research / reviewers in the wild / expert
Sherief Reda
dblp:11/6528
· DBLP profile ↗
110ranked-venue papers
10as first author
22since 2021 · last 2026
0000-0001-8232-4516ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 98 · 10 first-author · 15 since 2021Software engineering, systems software and programming languages · 18 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RobuMTL: Enhancing Multi-Task Learning Robustness Against Weather ConditionsabstractRobust Multi-Task Learning (MTL) is crucial for autonomous systems operating in real-world environments, where adverse weather conditions can severely degrade model performance and reliability. In this paper, we introduce RobuMTL, a novel architecture designed to adaptively address visual degradation by dynamically selecting taskspecific hierarchical Low-Rank Adaptation (LoRA) modules and a LoRA expert squad based on input perturbations in a mixture-of-experts fashion. Our framework enables adaptive specialization based on input characteristics, improving robustness across diverse real-world conditions. To validate our approach, we evaluated it on the PASCAL and NYUD-v2 datasets and compared it against single-task models, standard MTL baselines, and state-ofthe-art methods. On the PASCAL benchmark, RobuMTL delivers a +2.8% average relative improvement under single perturbations and up to +44.4% under mixed weather conditions compared to the MTL baseline. On NYUD-v2, RobuMTL achieves a +9.7% average relative improvement across tasks. The code will be available at GitHub1. Tasneem Shaffee, Sherief Reda |
WACV | 2 |
| 2025 | MetRex: A Benchmark for Verilog Code Metric Reasoning Using LLMsabstractLarge Language Models (LLMs) have been applied to various hardware design tasks, including Verilog code generation, EDA tool scripting, and RTL bug fixing. Despite this extensive exploration, LLMs are yet to be used for the task of post-synthesis metric reasoning and estimation of HDL designs. In this paper, we assess the ability of LLMs to reason about post-synthesis metrics of Verilog designs. We introduce MetRex, a large-scale dataset comprising 25,868 Verilog HDL designs and their corresponding post-synthesis metrics, namely area, delay, and static power. MetRex incorporates a Chain of Thought (CoT) template to enhance LLMs' reasoning about these metrics. Extensive experiments show that Supervised Fine-Tuning (SFT) boosts the LLM's reasoning capabilities on average by 37.0%, 25.3%, and 25.7% on the area, delay, and static power, respectively. While SFT improves performance on our benchmark, it remains far from achieving optimal results, especially on complex problems. Comparing to state-of-the-art regression models, our approach delivers accurate post-synthesis predictions for 17.4% more designs (within a 5% error margin), in addition to offering a 1.7x speedup by eliminating the need for pre-processing. This work lays the groundwork for advancing LLM-based Verilog code metric reasoning. Manar Abdelatty, Jingxiao Ma, Sherief Reda |
ASP-DAC | 3 |
| 2025 | FF-INT8: Efficient Forward-Forward DNN Training on Edge Devices with INT8 PrecisionabstractBackpropagation has been the cornerstone of neural network training for decades, yet its inefficiencies in time and energy consumption limit its suitability for resource-constrained edge devices. While low-precision neural network quantization has been extensively researched to speed up model inference, its application in training has been less explored. Recently, the Forward-Forward (FF) algorithm has emerged as a promising alternative to backpropagation, replacing the backward pass with an additional forward pass. By avoiding the need to store intermediate activations for backpropagation, FF can reduce memory footprint, making it well-suited for embedded devices. This paper presents an INT8 quantized training approach that leverages FF’s layer-by-layer strategy to stabilize gradient quantization. Furthermore, we propose a novel “look-ahead” scheme to address limitations of FF and improve model accuracy. Experiments conducted on NVIDIA Jetson Orin Nano board demonstrate 4.6% faster training, 8.3% energy savings, and $\mathbf{2 7. 0 \%}$ reduction in memory usage, while maintaining competitive accuracy compared to the state-of-the-art. Jingxiao Ma, Priyadarshini Panda, Sherief Reda |
DAC | 3 |
| 2025 | Fast Machine Learning Based Prediction for Temperature Simulation Using Compact ModelsabstractAs transistor densities increase, managing thermal challenges in 3D IC designs becomes more complex. Traditional methods like finite element methods and compact thermal models (CTMs) are computationally expensive, while existing machine learning (ML) models require large datasets and a long training time. To address these challenges with the ML models, we introduce a novel ML framework that integrates with CTMs to accelerate steady-state thermal simulations without needing large datasets. Our approach achieves up to 70 × speedup over state-of-the-art simulators, enabling real-time, high-resolution thermal simulations for 2D and 3D IC designs.11This research was partially funded by the NSF CCF 2131127 grant Mohammadamin Hajikhodaverdian, Sherief Reda, Ayse K. Coskun |
DATE | 2 |
| 2024 | MediatorDNN: Contention Mitigation for Co-Located DNN Inference JobsabstractWith the increase in computing power of cutting-edge hardware platforms, it is a common practice to run multiple jobs on a single machine for improved resource utilization and throughput. However, this leads to inevitable resource contention among co-located jobs, impacting their performance. The resource contention can worsen due to fluctuations in resource utilization of jobs caused by variations in their input workload. To tackle the co-location contention for DNN inference jobs, we propose MediatorDNN, which considers contention and resource utilization variation when co-locating DNN inference jobs. It profiles each DNN, monitors microarchitectural metrics such as memory bandwidth and cache access pattern, and high-level resource utilization like CPU utilization. Based on profiling results and leveraging Modern Portfolio Theory (MPT), MediatorDNN decides on the co-location of jobs. Experimental results with various DNNs on two hardware platforms show that MediatorDNN improves throughput by up to 108% (21 % on average) compared to an approach only considering contention and ignoring resource utilization variation. Seyed Morteza Nabavinejad, Sherief Reda, Tian Guo 0001 |
CLOUD | 2 |
| 2024 | PoliTune: Analyzing the Impact of Data Selection and Fine-Tuning on Economic and Political Biases in Large Language ModelsabstractIn an era where language models are increasingly integrated into decision-making and communication, understanding the biases within Large Language Models (LLMs) becomes imperative, especially when these models are applied in the economic and political domains. This work investigates the impact of fine-tuning and data selection on economic and political biases in LLMs. In this context, we introduce PoliTune, a fine-tuning methodology to explore the systematic aspects of aligning LLMs with specific ideologies, mindful of the biases that arise from their extensive training on diverse datasets. Distinct from earlier efforts that either focus on smaller models or entail resource-intensive pre-training, PoliTune employs Parameter-Efficient Fine-Tuning (PEFT) techniques, which allow for the alignment of LLMs with targeted ideologies by modifying a small subset of parameters. We introduce a systematic method for using the open-source LLM Llama3-70B for dataset selection, annotation, and synthesizing a preferences dataset for Direct Preference Optimization (DPO) to align the model with a given political ideology. We assess the effectiveness of PoliTune through both quantitative and qualitative evaluations of aligning open-source LLMs (Llama3-8B and Mistral-7B) to different ideologies. Our work analyzes the potential of embedding specific biases into LLMs and contributes to the dialogue on the ethical application of AI, highlighting the importance of deploying AI in a manner that aligns with societal values. Ahmed Agiza, Mohamed Mostagir, Sherief Reda |
AIES (1) | 3 |
| 2024 | MTLoRA: A Low-Rank Adaptation Approach for Efficient Multi-Task LearningabstractAdapting models pre-trained on large-scale datasets to a variety of downstream tasks is a common strategy in deep learning. Consequently, parameter-efficient fine-tuning methods have emerged as a promising way to adapt pretrained models to different tasks while training only a minimal number of parameters. While most of these methods are designed for single-task adaptation, parameter-efficient training in Multi-Task Learning (MTL) architectures is still unexplored. In this paper, we introduce MTLoRA, a novel framework for parameter-efficient training of MTL models. MTLoRA employs Task-Agnostic and Task-Specific Low-Rank Adaptation modules, which effectively disentangle the parameter space in MTL fine-tuning, thereby enabling the model to adeptly handle both task specialization and interaction within MTL contexts. We applied MTLoRA to hierarchical-transformer-based MTL architectures, adapting them to multiple downstream dense prediction tasks. Our extensive experiments on the PASCAL dataset show that MTLoRA achieves higher accuracy on downstream tasks compared to fully fine-tuning the MTL model while reducing the number of trainable parameters by 3.6×. Furthermore, MTLoRA establishes a Pareto-optimal trade-off between the number of trainable parameters and the accuracy of the downstream tasks, outperforming current state-of-the-art parameter-efficient training methods in both accuracy and efficiency. Our code is publicly available.11https://github.com/scale-lab/MTLoRA.git Ahmed Agiza, Marina Neseem, Sherief Reda |
CVPR | 3 |
| 2024 | PikeLPN: Mitigating Overlooked Inefficiencies of Low-Precision Neural NetworksabstractLow-precision quantization is recognized for its efficacy in neural network optimization. Our analysis reveals that non-quantized elementwise operations which are prevalent in layers such as parameterized activation functions, batch normalization, and quantization scaling dominate the inference cost of low-precision models. These non-quantized elementwise operations are commonly overlooked in SOTA efficiency metrics such as Arithmetic Computation Effort (ACE) [46]. In this paper, we propose ACEv2- an extended version of ACE which offers a better alignment with the inference cost of quantized models and their energy consumption on ML hardware. Moreover, we introduce PikeLPN11Pike is a slim fast fish, LPN stands for Low-Precision Network., a model that addresses these efficiency issues by applying quantization to both elementwise operations and multiply-accumulate operations. In particular, we present a novel quantization technique for batch normalization layers named QuantNorm which allows for quantizing the batch normalization parameters without compromising the model performance. Additionally, we propose applying Double Quantization where the quantization scaling parameters are quantized. Furthermore, we recognize and resolve the issue of distribution mismatch in Separable Convolution layers by introducing Distribution-Heterogeneous Quantization which enables quantizing them to low-precision. PikeLPN achieves Pareto-optimality in efficiency-accuracy trade-off with up to 3× efficiency improvement compared to SOTA low-precision models. Marina Neseem, Conor McCullough, Randy Hsin, Chas Leichner, Shan Li 0001, In Suk Chong, Andrew G. Howard, Lukasz Lew, Sherief Reda, Ville-Mikko Rautio, Daniele Moro |
CVPR | 9 |
| 2023 | RUCA: RUntime Configurable Approximate Circuits with Self-Correcting CapabilityabstractApproximate computing is an emerging computing paradigm that offers improved power consumption by relaxing the requirement for full accuracy. Since the requirements for accuracy may vary according to specific real-world applications, one trend of approximate computing is to design quality-configurable circuits, which are able to switch at runtime among different accuracy modes with different power and delay. In this paper, we present a novel framework RUCA which aims to synthesize runtime configurable approximate circuits based on arbitrary input circuits. By decomposing the truth table, our approach aims to approximate and separate the input circuit into multiple configuration blocks which support different accuracy levels, including a corrector circuit to restore full accuracy. Power gating is used to activate different blocks, such that the approximate circuit is able to operate at different accuracy-power configurations. To improve the scalability of our algorithm, we also provide a design space exploration scheme with circuit partitioning. We evaluate our methodology on a comprehensive set of benchmarks. For 3-level designs, RUCA saves power consumption by 43.71% within 2% error and by 30.15% within 1% error on average. Jingxiao Ma, Sherief Reda |
ASP-DAC | 2 |
| 2023 | GraPhSyM: Graph Physical Synthesis ModelabstractIn this work, we introduce GraPhSyM, a Graph Attention Network (GATv2) model for fast and accurate estimation of post-physical synthesis circuit delay and area metrics from pre-physical synthesis circuit netlists. Once trained, GraPhSyM provides accurate visibility of final design metrics to early EDA stages, such as logic synthesis, without running the slow physical synthesis flow, enabling global co-optimization across stages. Additionally, the swift and precise feedback provided by GraPhSym is instrumental for machine-learning-based EDA optimization frameworks. Given a gate-level netlist of a circuit represented as a graph, GraPhSyM utilizes graph structure, connectivity, and electrical property features to predict the impact of physical synthesis transformations such as buffer insertion and gate sizing. When trained on a dataset of 6000 prefix adder designs synthesized at an aggressive delay target, GraPhSyM can accurately predict the post-synthesis delay (98.3%) and area (96.1%) metrics of unseen adders with a fast 0.22s inference time. Furthermore, we illustrate the compositionality of GraPhSyM by employing the model trained on a fixed delay target to accurately anticipate post-synthesis metrics at a variety of unseen delay targets. Lastly, we report promising generalization capabilities of the GraPhSyM model when it is evaluated on circuits different from the adders it was exclusively trained on. The results show the potential for GraPhSyM to serve as a powerful tool for advanced optimization techniques and as an oracle for EDA machine learning frameworks. Ahmed Agiza, Rajarshi Roy 0003, Teodor-Dumitru Ene, Saad Godil, Sherief Reda, Bryan Catanzaro |
ICCAD | 5 |
| 2023 | WeNet: Configurable Neural Network with Dynamic Weight-Enabling for Efficient InferenceabstractDeep Neural Networks (DNN) are widely deployed in resource-limited edge devices. Due to the limitation of computational resources, it is important to meet the timing and energy constraints while maintaining a high level of accuracy. To deploy the same DNN model on different edge devices, one challenge is to train a dynamic neural network with the flexibility of balancing the trade-off between accuracy and efficiency at runtime. In this paper, we present a novel methodology, dynamic Weight-enabling Network (WeNet), where the weights of neural network can be dynamically enabled or disabled to switch between different sub-networks, so that we are able to balance the trade-off between inference time, energy consumption and model accuracy. We extend the methodology to convolutional layers using group convolution and channel shuffling. We also propose a design space exploration approach to search for the optimal sub-network for different scenarios. We thoroughly evaluate our methodology using a number of DNN architectures on different hardware platforms, showing that WeNet provides a large number of energy-efficient operation modes, 73.2 % of which provide better accuracy-efficiency trade-off compared to other methodologies. Jingxiao Ma, Sherief Reda |
ISLPED | 2 |
| 2022 | ARBench: Augmented Reality Benchmark For Mobile DevicesabstractThis paper takes an important step towards the improvement of the AR mobile experience by designing and developing ARBench, the first Augmented Reality (AR) benchmark for mobile devices. ARBench incorporates different AR workloads that stress multiple hardware units of the SoC (CPU, GPU, DSP, etc), and measures the individual score for each AR workload. The proposed benchmark suite is then used to evaluate the AR performance of various commercial mobile devices, and their ability to support various functions of AR workloads. Sofiane Chetoui, Rahul Shahi, Seif Abdelaziz, Abhinav Golas, Farrukh Hijaz, Sherief Reda |
ISPASS | 6 |
| 2022 | Alternating Blind Identification of Power Sources for Mobile SoCsabstractThe need for faster Systems on Chip (SoCs) has accelerated scaling trends, leading to a considerable power density increase and raising critical power and thermal challenges. The ability to measure power consumption of different hardware units is essential for the operation and improvement of mobile SoCs, as well as the enhancement of the power efficiency of the software that runs on them. SoCs are usually enabled with embedded thermal sensors to measure the temperature at the hardware unit level; however, they lack the ability to sense the power. In this paper we introduce an Alternating Blind Identification of Power sources (Alternating-BPI), a technique that accurately estimates the power consumption of individual SoC units without the use of any design based models. The proposed technique uses a novel approach to blindly identify the sources of power consumption, by relying only on the measurements from the embedded thermal sensors and the total power consumption. The accuracy and applicability of the proposed technique was verified using simulation and experimental data. Alternating-BPI is able to estimate the power at the SoC hardware unit level with up to 98.1% accuracy. Furthermore, we demonstrate the applicability of the proposed technique on a commercial SoC and provide a fine-grain analysis of the power profiles of CPU and GPU Apps, as well as Artificial Intelligence (AI), Virtual Reality (VR) and Augmented Reality (AR) Apps. Additionally, we demonstrate that the proposed technique could be used to estimate the power consumption per-process by relying on the estimated per-unit power numbers and per-unit hardware utilization numbers. The analysis provided by the proposed technique gives useful insights about the power efficiency of the different hardware units on a state-of-the-art commercial SoC. Sofiane Chetoui, Abhinav Golas, Farrukh Hijaz, Adel Belouchrani, Sherief Reda |
ICPE | 6 |
| 2022 | Characterizing and Optimizing EDA Flows for the CloudabstractDesign space exploration in logic synthesis and parameter tuning in physical design require a massive amount of compute resources in order to meet tapeout schedules. To address this need, cloud computing provides semiconductor and electronics companies with instant access to scalable compute resources. However, deploying electronic design automation (EDA) jobs on the cloud requires EDA teams to deeply understand the characteristics of their jobs in cloud environments. Unfortunately, there has been little to no public information on these characteristics. Thus, in this article, we first formulate the problem of moving EDA jobs to the cloud. To address the problem, we characterize the performance of four EDA main applications, namely: 1) synthesis; 2) placement; 3) routing; and 4) static timing analysis. We show that different EDA jobs require different compute configurations in order to achieve the best performance. Using observations from our characterization, we propose a novel model based on graph convolutional networks to predict the total runtime of a given stage on different configurations. Our model achieves a prediction accuracy of 87%. Furthermore, we present a new formulation for optimizing cloud deployments in order to reduce costs while meeting deadline constraints. We present a pseudopolynomial optimal solution using a multichoice knapsack mapping that reduces deployment costs by 35.29%, with minimal overhead to the total runtime. In addition, we describe a cloud-ready solution, called EDA analytics central, for the continuous optimization of a design across an EDA flow. We used this system in building our runtime prediction model. Abdelrahman Hosny, Sherief Reda |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | Approximate Logic Synthesis Using Boolean Matrix FactorizationabstractApproximate computing is an emerging computing paradigm offering benefits in hardware metrics, such as design area and power consumption, by relaxing the requirement for full accuracy. In circuit design, a major challenge is to synthesize approximate circuits automatically from input exact circuits requiring minimal expert input. In this work, we present a method for approximate logic synthesis based on the Boolean matrix factorization, where an arbitrary input circuit can be approximated in a controlled fashion. Our methodology enables automatic computation of the dominant elements,bases, of the truth table of the circuit, and later combines the bases to approximate the original truth table. Such compression can reduce the complexity of the hardware implementation significantly, while introducing variable degrees of inaccuracy. Furthermore, in our approach, the factorization algorithm can be fine tuned as required by the application, to effectively improve control over degree of approximation. In this work, we provide a unified approach enabling the factorization algorithm to utilize semiring algebra, field algebra, and a combination of both for truth table factorization. In addition, we provide an automatic circuit partitioning approach and a design space exploration heuristic to navigate the search space. We implement our methodology using a full stack of open-source tools, and thoroughly evaluate our methodology on a number of representative circuits showcasing the benefits of our proposed methodology for approximate logic synthesis. Finally, we compare our methodology against an existing library of approximate designs and demonstrate state-of-the-art performance. Jingxiao Ma, Soheil Hashemi, Sherief Reda |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | PACT: An Extensible Parallel Thermal Simulator for Emerging Integration and Cooling TechnologiesabstractThermal analysis is an essential step that enables co-design of the computing system (i.e., integrated circuits and computer architectures) with the cooling system (e.g., heat sink). Existing thermal simulation tools are limited by several major challenges that prevent them from providing fast solutions to large problem sizes that are necessary to conduct standard-cell level thermal analysis or to evaluate new technologies or large chips. To overcome these challenges, we introduce a SPICE-based parallel compact thermal simulator (PACT) that achieves fast and accurate, standard cell to architecture-level, steady-state, and transient parallel thermal simulations. PACT utilizes the advantages of multicore processing (OpenMPI) and includes several solvers to speed up both steady-state and transient simulations. PACT can be easily extended to model a variety of emerging integration and cooling technologies by simply modifying the thermal netlist. In addition, PACT can also be used with popular architecture-level performance and power simulators. In comparison to a state-of-the-art finite-element method (FEM)-based simulator (COMSOL), PACT has a maximum error of 2.77% and 3.28% for steady-state and transient thermal simulations, respectively. Compared to a popular compact thermal simulator, HotSpot, PACT demonstrates a speedup of up to$1.83\times $and$186\times $for steady-state and transient simulations, respectively. We also show the applicability and extensibility of PACT through modeling emerging integration and cooling technologies, such as monolithic 3-D integrated circuits and liquid cooling via microchannels, and full-system simulation integration on a 2.5-D system with silicon-photonic network-on-chips (PNoCs). Prachi Shukla, Sofiane Chetoui, Sean S. Nemtzow, Sherief Reda, Ayse K. Coskun |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2022 | Coordinated Batching and DVFS for DNN Inference on GPU AcceleratorsabstractEmploying hardware accelerators to improve the performance and energy-efficiency of DNN applications is on the rise. One challenge of using hardware accelerators, including the GPU-based ones, is that their performance is limited by internal and external factors, such as power caps. A common approach to meet the power cap constraint is using the Dynamic Voltage Frequency Scaling (DVFS) technique. However, the functionally of this technique is limited and platform-dependent. To tackle this challenge, we propose a new control knob, which is the size of input batches fed to the GPU accelerator in DNN inference applications. We first evaluate the impact of batch size on power consumption and performance of DNN inference. Then, we introduce the design and implementation of a fast and lightweight runtime system, called BatchDVFS. Dynamic batching is implemented inBatchDVFSto adaptively change the batch size, and hence, trade-off throughput with power consumption. It employs an approach based on binary search to find the proper batch size within a short period of time. Combining dynamic batching with the DVFS technique,BatchDVFScan control the power consumption in wider ranges, and hence, yield higher throughput in the presence of power caps. To find near-optimal solution for long-running jobs that can afford a relatively significant profiling overhead, compared withBatchDVFSoverhead, we also design an approach, called BOBD, that employs Bayesian Optimization to wisely explore the vast state space resulted by combination of the batch size and DVFS solutions. Conducting several experiments using a modern GPU and several DNN models and input datasets, we show that ourBatchDVFScan significantly surpass the techniques solely based on DVFS or batching, regarding throughput (up to 11.2x and 2.2x, respectively), while successfully meeting the power cap. Seyed Morteza Nabavinejad, Sherief Reda, Masoumeh Ebrahimi |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2021 | BatchSizer: Power-Performance Trade-off for DNN InferenceabstractGPU accelerators can deliver significant improvement for DNN processing; however, their performance is limited by internal and external parameters. A well-known parameter that restricts the performance of various computing platforms in real-world setups, including GPU accelerators, is the power cap imposed usually by an external power controller. A common approach to meet the power cap constraint is using the Dynamic Voltage Frequency Scaling (DVFS) technique. However, the functionally of this technique is limited and platform-dependent. To improve the performance of DNN inference on GPU accelerators, we propose a new control knob, which is the size of input batches fed to the GPU accelerator in DNN inference applications. After evaluating the impact of this control knob on power consumption and performance of GPU accelerators and DNN inference applications, we introduce the design and implementation of a fast and lightweight runtime system, called BatchSizer. This runtime system leverages the new control knob for managing the power consumption of GPU accelerators in the presence of the power cap. Conducting several experiments using a modern GPU and several DNN models and input datasets, we show that our BatchSizer can significantly surpass the conventional DVFS technique regarding performance (up to 29%), while successfully meeting the power cap. Seyed Morteza Nabavinejad, Sherief Reda, Masoumeh Ebrahimi |
ASP-DAC | 2 |
| 2021 | Workload- and User-aware Battery Lifetime Management for Mobile SoCsabstractMobile devices have become an essential part of daily life with the increased computing capabilities and features. For battery powered devices, the user experience depends on both quality-of-service (QoS) and battery lifetime. Previous works have been proposed to balance QoS and battery lifetime of mobile devices; however, they often consider only the CPU. Additionally, they fail in considering the user's desired battery lifetime while having a high QoS variation, which undermine the user satisfaction. In this work, we propose a CPU-GPU workload- and user-aware battery lifetime management technique for mobile devices using machine learning. Firstly, we design a workload-aware governor through an offline and an online analysis. A set of CPU and GPU performance counters is used during the offline analysis to identify a set of canonical phases (CP). In runtime, k-means is used to classify each sample of the performancecounters tooneof the predefined CP. Afterwards, we build a model that predicts the energy consumption given the user usage history. Finally, the energy model is used to find the optimal frequency settings for the CPU and GPU to provide the best QoS while meeting the target battery lifetime. The evaluation of the proposed work against state of the art techniques in a commercial smartphone, shows 15.8% and 9.4% performance improvement on the CPU and GPU, respectively. The proposed technique also shows 10× improvement in QoS variation, while meeting the desired battery lifetime. Sofiane Chetoui, Sherief Reda |
DATE | 2 |
| 2021 | Characterizing and Optimizing EDA Flows for the Cloud
Abdelrahman Hosny, Sherief Reda |
DATE | 2 |
| 2021 | AdaCon: Adaptive Context-Aware Object Detection for Resource-Constrained Embedded DevicesabstractConvolutional Neural Networks achieve state-of-the-art accuracy in object detection tasks. However, they have large computational and energy requirements that challenge their deployment on resource-constrained edge devices. Object detection takes an image as an input, and identifies the existing object classes as well as their locations in the image. In this paper, we leverage the prior knowledge about the probabilities that different object categories can occur jointly to increase the efficiency of object detection models. In particular, our technique clusters the object categories based on their spatial co-occurrence probability. We use those clusters to design an adaptive network. During runtime, a branch controller decides which part(s) of the network to execute based on the spatial context of the input frame. Our experiments using COCO dataset show that our adaptive object detection model achieves up to 45% reduction in the energy consumption, and up to 27% reduction in the latency, with a small loss in the average precision (AP) of object detection. Marina Neseem, Sherief Reda |
ICCAD | 2 |
| 2021 | Sparse Bitmap Compression for Memory-Efficient Training on the Edge
Abdelrahman Hosny, Marina Neseem, Sherief Reda |
SEC | 3 |
| 2020 | DRiLLS: Deep Reinforcement Learning for Logic SynthesisabstractLogic synthesis requires extensive tuning of the synthesis optimization flow where the quality of results (QoR) depends on the sequence of optimizations used. Efficient design space exploration is challenging due to the exponential number of possible optimization permutations. Therefore, automating the optimization process is necessary. In this work, we propose a novel reinforcement learning-based methodology that navigates the optimization space without human intervention. We demonstrate the training of an Advantage Actor Critic (A2C) agent that seeks to minimize area subject to a timing constraint. Using the proposed methodology, designs can be optimized autonomously with no-humans in-loop. Evaluation on the comprehensive EPFL benchmark suite shows that the agent outperforms existing exploration methodologies and improves QoRs by an average of 13%. Abdelrahman Hosny, Soheil Hashemi, Mohamed Shalan, Sherief Reda |
ASP-DAC | 4 |
| 2020 | ApproxDNN: Incentivizing DNN Approximation in CloudabstractService providers leverage discounted prices of reserved instances offered by cloud providers to amortize their operational costs. They reserve a certain number of instances to cover a significant portion of their computing resource requirements, and further employ on-demand instances to cover remaining requirements not satisfied by the reserved instances. Because of the higher price of on-demand instances, service providers seek to lower their usage to minimize operational costs. In this work, we propose ApproxDNN approach for Machine Learning as a Service to reduce operational costs of service providers by incentivizing approximate results, based on the capabilities of cutting-edge GPUs and a discounted pricing model. When the deadlines of jobs submitted by users are very tight, a service provider might not be able to execute all of them on reserved instances under the default precision. In such cases, Ap- proxDNN leverages the reduced-precision instructions to reduce the execution time of the jobs with slight reduction in their final accuracy, and consequently, to minimize the employment of on- demand instances. To incentivize users to accept the approximate results of reduced-precision instructions, ApproxDNN offers them a discounted price for the service based on a newly designed pricing model. Our proposed pricing model of ApproxDNN guarantees lower or equal cost for service providers compared to the conventional method that solely depends on employment of on-demand instances in case of the reserved instance shortage. We employ real-world traces to conduct an extensive set of experiments and evaluate the performance of our proposed approach. The results show that ApproxDNN reduces the cost of service providers by 18%, while never exceeding the cost of the conventional method and slightly affecting the accuracy by 0.14%. Seyed Morteza Nabavinejad, Lena Mashayekhy, Sherief Reda |
CCGRID | 3 |
| 2020 | AdaSense: Adaptive Low-Power Sensing and Activity Recognition for Wearable DevicesabstractWearable devices have strict power and memory limitations. As a result, there is a need to optimize the power consumption on those devices without sacrificing the accuracy. This paper presents AdaSense: a sensing, feature extraction and classification co-optimized framework for Human Activity Recognition. The proposed techniques reduce the power consumption by dynamically switching among different sensor configurations as a function of the user activity. The framework selects configurations that represent the pareto-frontier of the accuracy and energy trade-off. AdaSense also uses low-overhead processing and classification methodologies. The introduced approach achieves 69% reduction in the power consumption of the sensor with less than 1.5% decrease in the activity recognition accuracy. Marina Neseem, Jon Nelson, Sherief Reda |
DAC | 3 |
| 2020 | A Learning-Based Thermal Simulation Framework for Emerging Two-Phase Cooling TechnologiesabstractFuture high-performance chips will require new cooling technologies that can extract heat efficiently. Two-phase cooling is a promising processor cooling solution owing to its high heat transfer rate and potential benefits in cooling power. Two-phase cooling mechanisms, including microchannel-based two-phase cooling or two-phase vapor chambers (VCs), are typically modeled by computing the temperature-dependent heat transfer coefficient (HTC) of the evaporator or coolant using an iterative simulation framework. Precomputed HTC correlations are specific to a given cooling system design and cannot be applied to even the same cooling technology with different cooling parameters (such as different geometries). Another challenge is that HTC correlations are typically calculated with computational fluid dynamics (CFD) tools, which induce long design and simulation times. This paper introduces a learning-based temperature-dependent HTC simulation framework that is used to model a two-phase cooling solution with a wide range of cooling design parameters. In particular, the proposed framework includes a compact thermal model (CTM) of two-phase VCs with hybrid wick evaporators (of nanoporous membrane and microchannels). We build a new simulation tool to integrate the proposed simulation framework and CTM. We validate the proposed simulation framework as well as the new CTM through comparisons against a CFD model. Our simulation framework and CTM achieve a speedup of 21 × with an average error of 0.98° C (and a maximum error of 2.59° C). We design an optimization flow for hybrid wicks to select the most beneficial hybrid wick geometries. Our flow is capable of finding a geometry- coolant combination that results in a lower (or similar) maximum chip temperature compared to that of the best coolant-geometry pair selected by grid search, while providing a speedup of 9.4 x. Geoffrey Vaartstra, Prachi Shukla, Zhengmao Lu, Evelyn Wang, Sherief Reda, Ayse K. Coskun |
DATE | 6 |
| 2020 | Temperature and Supply Voltage Monitoring with Current-mode Relaxation OscillatorsabstractThis paper presents a new family of temperature and supply voltage sensors based on an improved current-mode relaxation oscillator. The proposed sensor circuit overcomes the finite input impedance of earlier current-mode relaxation oscillators, and it is paired with a novel opamp-free bandgap current reference. Two examples are presented: a VDD- controlled oscillator, and a temperature-controlled oscillator. The temperature-controlled oscillator operates at 11.3 MHz and consumes 11.14 uW while reaching a temperature nonlinearity error less than +0.85/-0.94°C, a resolution figure of merit of 11.5 pJ·K2, and an efficiency of 240 pJ/conv. The VDD-controlled oscillator has a nonlinearity error less than +28.7/-30.0mV from 1.2-1.8 V after two-point calibration. Shanshan Dai, Caleb R. Tulloss, Xiaoyu Lian, Kangping Hu, Sherief Reda, Jacob K. Rosenstein |
VLSI-SOC | 5 |
| 2020 | Simultaneous Estimation of Temperature and Voltage from Digital Delay DiversityabstractSemiconductor devices are fundamentally sensitive to process variation, supply voltage, and temperature, which impacts both the design and performance of modern integrated circuits. On-chip sensors are typically designed and calibrated to measure temperature and voltage independently, but designing independent sensors has implicit power and area costs. Here we propose a new simultaneous estimation of temperature and supply voltage using an ensemble of digital logic delays and nonlinear regression models. This approach is inherently digital and compatible with modern digital design flows. Using simulations in commercial 12nm FinFET and 65nm bulk CMOS processes, we show the feasibility and efficiency of this approach, and consider its benefits and potential limitations. Xiaoyu Lian, Sherief Reda, Jacob K. Rosenstein |
VLSI-SOC | 2 |
| 2020 | Dual-precision fixed-point arithmetic for low-power ray-triangle intersections
Krishna Rajan, Soheil Hashemi, Ulya R. Karpuzcu, Michael C. Doggett, Sherief Reda |
Comput. Graph. | 5 |
| 2020 | A Resource-Efficient Embedded Iris Recognition System Using Fully Convolutional NetworksabstractApplications of fully convolutional networks (FCN) in iris segmentation have shown promising advances. For mobile and embedded systems, a significant challenge is that the proposed FCN architectures are extremely computationally demanding. In this article, we propose a resource-efficient, end-to-end iris recognition flow, which consists of FCN-based segmentation and a contour fitting module, followed by Daugman normalization and encoding. To attain accurate and efficient FCN models, we propose a three-step SW/HW co-design methodology consisting of FCN architectural exploration, precision quantization, and hardware acceleration. In our exploration, we propose multiple FCN models, and in comparison to previous works, our best-performing model requires 50× fewer floating-point operations per inference while achieving a new state-of-the-art segmentation accuracy. Next, we select the most efficient set of models and further reduce their computational complexity through weights and activations quantization using an 8-bit dynamic fixed-point format. Each model is then incorporated into an end-to-end flow for true recognition performance evaluation. A few of our end-to-end pipelines outperform the previous state of the art on two datasets evaluated. Finally, we propose a novel dynamic fixed-point accelerator and fully demonstrate the SW/HW co-design realization of our flow on an embedded FPGA platform. In comparison with the embedded CPU, our hardware acceleration achieves up to 8.3× speedup for the overall pipeline while using less than 15% of the available FPGA resources. We also provide comparisons between the FPGA system and an embedded GPU showing different benefits and drawbacks for the two platforms. Hokchhay Tann, Sherief Reda |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2020 | Approximate Logic Synthesis: A SurveyabstractApproximate computing is an emerging paradigm that, by relaxing the requirement for full accuracy, offers benefits in terms of design area and power consumption. This paradigm is particularly attractive in applications where the underlying computation has inherent resilience to small errors. Such applications are abundant in many domains, including machine learning, computer vision, and signal processing. In circuit design, a major challenge is the capability to synthesize the approximate circuits automatically without manually relying on the expertise of designers. In this work, we review methods devised to synthesize approximate circuits, given their exact functionality and an approximability threshold. We summarize strategies for evaluating the error that circuit simplification can induce on the output, which guides synthesis techniques in choosing the circuit transformations that lead to the largest benefit for a given amount of induced error. We then review circuit simplification methods that operate at the gate or Boolean level, including those that leverage classical Boolean synthesis techniques to realize the approximations. We also summarize strategies that take high-level descriptions, such as C or behavioral Verilog, and synthesize approximate circuits from these descriptions. Ilaria Scarabottolo, Giovanni Ansaloni, George A. Constantinides, Laura Pozzi 0001, Sherief Reda |
Proc. IEEE | 5 |
| 2020 | LoCool: Fighting Hot Spots Locally for Improving System Energy EfficiencyabstractElevated on-chip temperatures significantly degrade performance, energy-efficiency, and lifetime of processors. The cooling system for a chip is typically designed to remove the worst-case heat generated per unit area. Cooling demand, however, spatially and temporally varies across a chip as hot spots occur on different locations with different intensities. Thus, designing a homogeneous cooling system for a chip can be inefficient. Recently, hybrid cooling strategies, such as integrating thermoelectric coolers (TECs) with microchannel liquid cooling, have been explored for hot spot mitigation. The efficiency of such a cooling system strongly depends on the operating point of each cooling method, as well as the locations and intensities of the hot spots. To this end, we first devise a compact thermal modeling method for the design and evaluation of hybrid cooling systems in a fast and accurate way. The proposed model provides up to four orders of magnitude speedup in simulation time compared to COMSOL multiphysics simulations with less than 2.9 °C average temperature error. Leveraging our fast model, we develop LoCool, a hybrid cooling optimization method, which jointly determines the most energy-efficient cooling settings for a given chip power distribution and temperature constraint. LoCool determines the liquid flow rate and the input current for each TEC depending on the cooling requirements for individual hot spots as well as for the background heat. Experimental evaluation shows up to 40% cooling energy savings compared to designing homogeneous cooling systems under the same thermal constraints. Fulya Kaplan, Mostafa Said, Sherief Reda, Ayse K. Coskun |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | Toward an Open-Source Digital Flow: First Learnings from the OpenROAD ProjectabstractWe describe the planned Alpha release of OpenROAD, an open-source end-to-end silicon compiler. OpenROAD will help realize the goal of "democratization of hardware design", by reducing cost, expertise, schedule and risk barriers that confront system designers today. The development of open-source, self-driving design tools is in and of itself a "moon shot" with numerous technical and cultural challenges. The open-source flow incorporates a compatible open-source set of tools that span logic synthesis, floorplanning, placement, clock tree synthesis, global routing and detailed routing. The flow also incorporates analysis and support tools for static timing analysis, parasitic extraction, power integrity analysis, and cloud deployment. We also note several observed challenges, or "lessons learned", with respect to development of open-source EDA tools and flows. Tutu Ajayi, Vidya A. Chhabria, Mateus Fogaça, Soheil Hashemi, Abdelrahman Hosny, Andrew B. Kahng, Jeongsup Lee, Uday Mallappa, Marina Neseem, Geraldo Pradipta, Sherief Reda, Mehdi Saligane, Sachin S. Sapatnekar, Carl Sechen, Mohamed Shalan, William Swartz, Lutong Wang, Zhehong Wang, Mingyu Woo, Bangqi Xu |
DAC | 12 |
| 2019 | Generalized Matrix Factorization Techniques for Approximate Logic SynthesisabstractApproximate computing is an emerging computing paradigm, where computing accuracy is relaxed for improvements in hardware metrics, such as design area and power profile. In circuit design, a major challenge is to synthesize approximate circuits automatically from input exact circuits. In this work, we extend our previous work, BLASYS, for approximate logic synthesis based on matrix factorization, where an arbitrary input circuit can be approximated in a controlled fashion. Whereas our previous approach uses a semi-ring algebra for factorization, this work generalizes matrix-based circuit factorization to include both semi-ring and field algebra implementations. We also propose a new method for truth table folding to improve the factorization quality. These new approaches significantly widen the design space of possible approximate circuits, effectively offering improved trade-offs in terms of quality, area and power consumption. We evaluate our methodology on a number of representative circuits showcasing the benefits of our proposed methodology for approximate logic synthesis. Soheil Hashemi, Sherief Reda |
DATE | 2 |
| 2019 | Modeling and Optimization of Chip Cooling with Two-Phase Vapor ChambersabstractUltra-high power densities that are expected in future processors cannot be efficiently mitigated by conventional cooling solutions. Using two-phase vapor chambers (VCs) with micropillar wick evaporators is an emerging cooling technique that can effectively remove high heat fluxes through the evaporation process of a coolant. Two-phase VCs with micropillar wicks offer high cooling efficiency by leveraging a capillary-driven flow, where the coolant is passively driven by the wicking structure that eliminates the need for an external pump. Thermal models for such emerging cooling technologies are essential to evaluate their impact on future processors. Existing thermal models for two-phase VCs use computational fluid dynamics (CFD) modules, which incur long design and simulation times. This paper presents a fast and accurate compact thermal model for two-phase VCs with micropillar wicks. Our model achieves a maximum error of 1.25°C with a speedup of 214x in comparison to a CFD model. Using our proposed thermal model, we build an optimization flow that selects the best cooling solution and its cooling parameters to minimize the cooling power under a temperature constraint for a given processor and power profile. We then demonstrate our optimization flow on different chip sizes and hot spot distributions to choose the optimal cooling technique among VCs, microchannel-based two-phase cooling, liquid cooling via microchannels, and a hybrid cooling technique with thermoelectric coolers and liquid cooling with microchannels. Geoffrey Vaartstra, Prachi Shukla, Sherief Reda, Evelyn Wang, Ayse K. Coskun |
ISLPED | 4 |
| 2018 | BLASYS: approximate logic synthesis using boolean matrix factorizationabstractApproximate computing is an emerging paradigm where design accuracy can be traded off for benefits in design metrics such as design area, power consumption or circuit complexity. In this work, we present a novel paradigm to synthesize approximate circuits using Boolean matrix factorization (BMF). In our methodology the truth table of a sub-circuit of the design is approximated using BMF to a controllable approximation degree, and the results of the factorization are used to synthesize a less complex subcircuit. To scale our technique to large circuits, we devise a circuit decomposition method and a subcircuit design-space exploration technique to identify the best order for subcircuit approximations. Our method leads to a smooth trade-off between accuracy and full circuit complexity as measured by design area and power consumption. Using an industrial strength design flow, we extensively evaluate our methodology on a number of testcases, where we demonstrate that the proposed methodology can achieve up to 63% in power savings, while introducing an average relative error of 5%. We also compare our work to previous works in Boolean circuit synthesis and demonstrate significant improvements in design metrics for same accuracy targets. Soheil Hashemi, Hokchhay Tann, Sherief Reda |
DAC | 3 |
| 2018 | Approximate computing for biometrie security systems: A case study on iris scanningabstractExploiting the error resilience of emerging data-rich applications, approximate computing promotes the introduction of small amount of inaccuracy into computing systems to achieve significant reduction in computing resources such as power, design area, runtime or energy. Successful applications for approximate computing have been demonstrated in the areas of machine learning, image processing and computer vision. In this paper we make the case for a new direction for approximate computing in the field of biometric security with a comprehensive case study of iris scanning. We devise an end-to-end flow from an input camera to the final iris encoding that produces sufficiently accurate final results despite relying on intermediate approximate computational steps. Unlike previous methods which evaluated approximate computing techniques on individual algorithms, our flow consists of a complex SW/HW pipeline of four major algorithms that eventually compute the iris encoding from input live camera feeds. In our flow, we identify overall eight approximation knobs at both the algorithmic and hardware levels to trade-off accuracy with runtime. To identify the optimal values for these knobs, we devise a novel design space exploration technique based on reinforcement learning with a recurrent neural network agent. Finally, we fully implement and test our proposed methodologies using both benchmark dataset images and live images from a camera using an FPGA-based SoC. We show that we are able to reduce the runtime of the system by 48 χ on top of an already HW accelerated design, while meeting industry-standard accuracy requirements for iris scanning systems. Soheil Hashemi, Hokchhay Tann, Francesco Buttafuoco, Sherief Reda |
DATE | 4 |
| 2018 | QoR-aware power capping for approximate big data processingabstractTo limit the peak power consumption of a cluster, a centralized power capping system typically assigns power caps to the individual servers, which are then enforced using local capping controllers. Consequently, the performance and throughput of the servers are affected, and the runtime of jobs is extended as a result. We observe that servers in big data processing clusters often execute big data applications that have different tolerance for approximate results. To mitigate the impact of power capping, we propose a new power-Capping aware resource manager for Approximate Big data processing (CAB) that takes into consideration the minimum Quality-of-Result (QoR) of the jobs. We use industry-standard feedback power capping controllers to enforce a power cap quickly, while, simultaneously modifying the resource allocations to various jobs based on their progress rate, target minimum QoR, and the power cap such that the impact of capping on runtime is minimized. Based on the applied cap and the progress rates of jobs, CAB dynamically allocates the computing resources (i.e., number of cores and memory) to the jobs to mitigate the impact of capping on the finish time. We implement CAB in Hadoop-2.7.3 and evaluate its improvement over other methods on a state-of-the-art 28-core Xeon server. We demonstrate that CAB minimizes the impact of power capping on runtime by up to 39.4% while meeting the minimum QoR constraints. Seyed Morteza Nabavinejad, Xin Zhan, Maziar Goudarzi, Sherief Reda |
DATE | 5 |
| 2018 | Computing with Chemicals: Perceptrons Using Mixtures of Small MoleculesabstractComputation that can exploit the Avogadrian numbers of molecules in heterogeneous solutions, and the even larger number of potential interactions among these molecules, is a tantalizing dream. However, the lack of precise specificity/control of chemical interactions can be at odds with the dream. In this paper, we show how relatively simple chemistry can be used to produce a ubiquitous computational primitive (the multiply-accumulate or MAC operation) that forms the basis for a single-layer neural network called a perceptron. A chemical perceptron can be realized using distinct mixtures as inputs and different reagents as operations to produce the results of the perceptron MAC operation, that can be read out perhaps using simple indicators such as pH or fluorescence. With a moderately large chemical library, the number of potential inputs can be Avogadrian so that reagent addition implicitly performs a concomitantly large number of MAC operations in parallel. Christopher Rose, Sherief Reda, Brenda M. Rubenstein, Jacob K. Rosenstein |
ISIT | 2 |
| 2017 | Understanding the Role of GPGPU-Accelerated SoC-Based ARM ClustersabstractThe last few years saw the emergence of 64-bit ARM SoCs targeted for mobile systems and servers. Mobile-class SoCs rely on the heterogeneous integration of a mix of CPU cores, GPGPU cores, and accelerators, whereas server-class SoCs instead rely on integrating a larger number of CPU cores with no GPGPU support and a number of network accelerators. Previous works, such as the Mont-Blanc project, built their prototype ARM cluster out of mobile-class SoCs and compared their work against x86 solutions. These works mainly focused on the CPU performance. In this paper, we propose a novel ARM-based cluster organization that exploits faster network connectivity and GPGPU acceleration to improve the performance and energy efficiency of the cluster. Our custom cluster, based on Nvidia Jetson TX1 boards, is equipped with 10Gb network interface cards and enables us to study the characteristics, scalability challenges, and programming models of GPGPU-accelerated workloads. We also develop an extension to the Roofline model to establish a visually intuitive performance model for the proposed cluster organization. We compare the GPGPU performance of our cluster with discrete GPGPUs. We demonstrate that our cluster improves both the performance and energy efficiency of workloads that scale well and can leverage the better CPU-GPGPU balance of our cluster. We contrast the CPU performance of our cluster with ARM-based servers that use many CPU cores. Our results show the poor performance of the branch predictor and L2 cache are the bottleneck of server-class ARM SoCs. Furthermore, we elucidate the impact of using 10Gb connectivity with mobile systems instead of traditional, 1Gb connectivity. Tyler Fox, Sherief Reda |
CLUSTER | 3 |
| 2017 | Hardware-Software Codesign of Accurate, Multiplier-free Deep Neural NetworksabstractWhile Deep Neural Networks (DNNs) push the state-of-the-art in many machine learning applications, they often require millions of expensive floating-point operations for each input classification. This computation overhead limits the applicability of DNNs to low-power, embedded platforms and incurs high cost in data centers. This motivates recent interests in designing low-power, low-latency DNNs based on fixed-point, ternary, or even binary data precision. While recent works in this area offer promising results, they often lead to large accuracy drops when compared to the floating-point networks. We propose a novel approach to map floating-point based DNNs to 8-bit dynamic fixed-point networks with integer power-of-two weights with no change in network architecture. Our dynamic fixed-point DNNs allow different radix points between layers. During inference, power-of-two weights allow multiplications to be replaced with arithmetic shifts, while the 8-bit fixed-point representation simplifies both the buffer and adder design. In addition, we propose a hardware accelerator design to achieve low-power, low-latency inference with insignificant degradation in accuracy. Using our custom accelerator design with the CIFAR-10 and ImageNet datasets, we show that our method achieves significant power and energy savings while increasing the classification accuracy. Hokchhay Tann, Soheil Hashemi, R. Iris Bahar, Sherief Reda |
DAC | 4 |
| 2017 | Understanding the impact of precision quantization on the accuracy and energy of neural networksabstractDeep neural networks are gaining in popularity as they are used to generate state-of-the-art results for a variety of computer vision and machine learning applications. At the same time, these networks have grown in depth and complexity in order to solve harder problems. Given the limitations in power budgets dedicated to these networks, the importance of low-power, low-memory solutions has been stressed in recent years. While a large number of dedicated hardware using different precisions has recently been proposed, there exists no comprehensive study of different bit precisions and arithmetic in both inputs and network parameters. In this work, we address this issue and perform a study of different bit-precisions in neural networks (from floating-point to fixed-point, powers of two, and binary). In our evaluation, we consider and analyze the effect of precision scaling on both network accuracy and hardware metrics including memory footprint, power and energy consumption, and design area. We also investigate training-time methodologies to compensate for the reduction in accuracy due to limited bit precision and demonstrate that in most cases, precision scaling can deliver significant benefits in design metrics at the cost of very modest decreases in network accuracy. In addition, we propose that a small portion of the benefits achieved when using lower precisions can be forfeited to increase the network size and therefore the accuracy. We evaluate our experiments, using three well-recognized networks and datasets to show its generality. We investigate the trade-offs and highlight the benefits of using lower precisions in terms of energy and memory footprint. Soheil Hashemi, Nicholas Anthony, Hokchhay Tann, R. Iris Bahar, Sherief Reda |
DATE | 5 |
| 2017 | Blind identification of power sources in processorsabstractThe ability to measure power consumption is at the heart of power and thermal management techniques. Modern processors are equipped with hardware monitoring mechanisms that can measure total power. However, this lumped measurement is not sufficient if there is a need to execute fine-grain thermal and power management techniques. This paper proposes a new direction for identifying the fine-grain sources of power consumption in many-core processors. For the first time, we show that it is possible to simultaneously identify both the power consumption of different cores and the thermal model of the chip from just the measurements of the thermal sensors and the total power consumption measurement. Our identification technique is blind as it does not require design knowledge of the thermal model to identify the power sources. Furthermore, our technique makes no use of the performance counters, which reduces its overhead, and works seamlessly with dynamic voltage and frequency scaling. We implement our technique on a real multi-core CPU-GPU processor-based system, and we show the ability to identify the runtime power consumption of the individual cores using just the total power measurement and the measurements of the thermal sensors under different workloads. We also verify the superior accuracy of our approach using results from a controlled simulation environment. Sherief Reda, Adel Belouchrani |
DATE | 1 |
| 2017 | Fast Decentralized Power Capping for Server ClustersabstractPower capping is a mechanism to ensure that the power consumption of clusters does not exceed the provisioned resources. A fast power capping method allows for a safe over-subscription of the rated power distribution devices, provides equipment protection, and enables large clusters to participate in demand-response programs. However, current methods have a slow response time with a large actuation latency when applied across a large number of servers as they rely on hierarchical management systems. We propose a fast decentralized power capping (DPC) technique that reduces the actuation latency by localizing power management at each server. The DPC method is based on a maximum throughput optimization formulation that takes into account the workloads priorities as well as the capacity of circuit breakers. Therefore, DPC significantly improves the cluster performance compared to alternative heuristics. We implement the proposed decentralized power management scheme on a real computing cluster. Compared to state-of-the-art hierarchical methods, DPC reduces the actuation latency by 72% up to 86% depending on the cluster size. In addition, DPC improves the system throughput performance by 16%, while using only 0.02% of the available network bandwidth. We describe how to minimize the overhead of each local DPC agent to a negligible amount. We also quantify the traffic and fault resilience of our decentralized power capping approach. Masoud Badiei, Xin Zhan, Na Li 0002, Sherief Reda |
HPCA | 5 |
| 2017 | LACore: A Supercomputing-Like Linear Algebra Accelerator for SoC-Based DesignsabstractLinear algebra operations are at the heart of scientific computing solvers, machine learning and artificial intelligence. In this paper, LACore, a novel, programmable accelerator architecture for general-purpose linear algebra applications, is presented. LACore enables many of the architectural features typically available in custom supercomputing machines in an accelerator form factor that can be deployed in System-On-a-Chip (SoC) based designs. LACore has several architectural features including heterogeneous data-streaming LAMemUnits, a configurable systolic datapath that supports scalar, vector and multi-stream output modes, and a decoupled architecture that overlap memory transfer and execution. To evaluate LACore, we implemented its architecture as an extension to the RISC-V ISA in the gem5 cycle-accurate simulator. The LACore ISA was implemented in gcc, and a C-programming software framework, the LACoreAPI, has been developed for high-level programming of the LACore. Using the HPCC benchmark suite, we compare our LACore architecture against three other platforms: an in-order RISC-V CPU, a superscalar x86 CPU with SSE2, and a scaled NVIDIA Fermi GPU. The LACore outperforms the superscalar x86 processor in the benchmark suite by an average of 3.43x, and outperforms the scaled Fermi GPU by an average of 12.04x, within the same or less design area. Samuel Steffl, Sherief Reda |
ICCD | 2 |
| 2016 | DiBA: Distributed Power Budget Allocation for Large-Scale Computing ClustersabstractPower management has become a central issue inlarge-scale computing clusters where a considerable amount ofenergy is consumed and a large operational cost is incurredannually. Traditional power management techniques have a centralizeddesign that creates challenges for scalability of computingclusters. In this work, we develop a framework for distributedpower budget allocation that maximizes the utility of computingnodes subject to a total power budget constraint. To eliminate the role of central coordinator in the primaldualtechnique, we propose a distributed power budget allocationalgorithm (DiBA) which maximizes the combined performanceof a cluster subject to a power budget constraint in a distributedfashion. Specifically, DiBA is a consensus-based algorithm inwhich each server determines its optimal power consumptionlocally by communicating its state with neighbors (connectednodes) in a cluster. We characterize a synchronous primal-dualtechnique to obtain a benchmark for comparison with thedistributed algorithm that we propose. We demonstrate numericallythat DiBA is a scalable algorithm that outperforms theconventional primal-dual method on large scale clusters in termsof convergence time. Further, DiBA eliminates the communicationbottleneck in the primal-dual method. We thoroughly evaluatethe characteristics of DiBA through simulations of large-scaleclusters. Furthermore, we provide results from a proof-of-conceptimplementation on a real experimental cluster. Masoud Badiei, Xin Zhan, Sherief Reda, Na Li 0002 |
CCGrid | 4 |
| 2016 | Creating Soft Heterogeneity in Clusters Through Firmware Re-configurationabstractCustomizing server hardware to adapt to its workload has the potential to improve both runtime and energy efficiency. In a cluster that caters to diverse workloads, employing servers with customized hardware components leads to heterogeneity, which is not scalable. In this paper, we seek to create soft heterogeneity from existing servers with homogenous hardware components through customizing the firmware configuration. We demonstrate that firmware configurations have a large impact on runtime, power, and energy efficiency of workloads. Since finding the firmware configuration that minimizes runtime and/or energy efficiency grows exponentially as a function of the number of firmware settings, we propose a methodology called FXplore that helps complete the exploration with a quadratic time complexity. Furthermore, FXplore enables system administrators to manage the degree of the heterogeneity by deriving firmware configurations for sub-clusters that can cater to multiple workloads with similar characteristics. Thus, during online operation, incoming workloads to the cluster can be mapped to appropriate sub-clusters with pre-configured firmware settings. FXplore also finds the best firmware settings in case of co-runners on the same server. We validate our methodology on a fully-instrumented cluster under a large range of parallel workloads that are representative of both high-performance compute clusters and datacenters. Compared to enabling all firmware options, our method improves average runtime and energy consumption by 11% and 15%, respectively. Xin Zhan, Mohammed Shoaib, Sherief Reda |
CCGrid | 3 |
| 2016 | A low-power dynamic divider for approximate applicationsabstractIn this work, a low-power, low-error divider design is proposed that can achieve significant power and area savings, while introducing insignificant inaccuracies to the output. The design of our divider is highly scalable, offering a wide range of power and inaccuracy trade-offs based on the application requirements. Furthermore, the proposed divider has a lower delay compared to the accurate design, enabling its use on the critical path. We theoretically analyze the error of our design as a function of its configuration, and we thoroughly evaluate the error and power characteristics of our divider in a standalone case and demonstrate that the proposed design can achieve up to 70% in power savings, while introducing an mean average absolute error of only 3.08%. We also implement three image-processing applications in hardware using our divider and conclude that use of the proposed divider will not perceptibly impact their quality-of-service while achieving power benefits of up to 75%. Soheil Hashemi, R. Iris Bahar, Sherief Reda |
DAC | 3 |
| 2016 | Hardware acceleration of feature detection and description algorithms on low-power embedded platformsabstractImage features are broadly used in embedded computer vision applications, from object detection and tracking to motion estimation and 3D reconstruction. Efficient feature extraction and description are crucial due to the real-time requirements of such applications over a constant stream of input data. High-speed computation typically comes at the cost of high power dissipation, yet embedded systems are often highly power constrained, making discovery of power-aware solutions especially critical for these systems. In this paper, we present a power and performance evaluation of three low cost feature detection and description algorithms implemented on various embedded systems (embedded CPUs, GPUs and FPGAs). We show that FPGAs in particular offer attractive solutions for both performance and power and describe several design techniques utilized to accelerate feature extraction and description algorithms on low-cost Zynq SoC FPGAs. Onur Ulusel, Christopher B. Picardo, Christopher B. Harris, Sherief Reda, R. Iris Bahar |
FPL | 4 |
| 2015 | DRUM: A Dynamic Range Unbiased Multiplier for Approximate ApplicationsabstractMany applications for signal processing, computer vision and machine learning show an inherent tolerance to some computational error. This error resilience can be exploited to trade off accuracy for savings in power consumption and design area. Since multiplication is an essential arithmetic operation for these applications, in this paper we focus specifically on this operation and propose a novel approximate multiplier with a dynamic range selection scheme. We design the multiplier to have an unbiased error distribution, which leads to lower computational errors in real applications because errors cancel each other out, rather than accumulate, as the multiplier is used repeatedly for a computation. Our approximate multiplier design is also scalable, enabling designers to parameterize it depending on their accuracy and power targets. Furthermore, our multiplier benefits from a reduction in propagation delay, which enables its use on the critical path. We theoretically analyze the error of our design as a function of its parameters and evaluate its performance for a number of applications in image processing, and machine classification. We demonstrate that our design can achieve power savings of 54% - 80%, while introducing bounded errors with a Gaussian distribution with near-zero average and standard deviations of 0.45% - 3.61%. We also report power savings of up to 58% when using the proposed design in applications. We show that our design significantly outperforms other approximate multipliers recently proposed in the literature. Soheil Hashemi, R. Iris Bahar, Sherief Reda |
ICCAD | 3 |
| 2015 | Making sense of thermoelectrics for processor thermal management and energy harvestingabstractA thermoelectric (TE) device can be used as a heat pump that consumes electric power to cool a processor chip, or it can be used as a heat engine that generates electricity from the heat dissipated during processor operation. To better understand the use of TE devices, we develop a fully instrumented processor-based system with controllable TE devices. We first examine the use of TE devices for energy harvesting. We identify a pitfall in previous works that can lead to wrong conclusions for TEG use by demonstrating that TEGs increase the processor's leakage power which offsets their harvested power. For thermoelectric cooling (TEC), we elucidate the intricate relationships between the processor power, thermoelectric power, and fan power. We propose a dynamic thermal management scheme (DTM) that maximizes performance under thermal constraints and given total power budgets by controlling the processor's dynamic frequency and voltage scaling (DVFS), TEC current, and fan speed. For the evaluated thermal constraints, our results demonstrate good improvements to performance at the cost of additional cooling power compared to standard DVFS+fan DTM techniques. Sriram Jayakumar, Sherief Reda |
ISLPED | 2 |
| 2015 | Power Budgeting Techniques for Data CentersabstractThe development of cloud computing and data science result in rapid increases of number and scale of data centers. Because of cost and sustainability concerns, energy efficiency has been a major goal for data center architects. Focusing on reducing the cooling power and making full use of available computing power, power budgeting is an increasingly important requirement for data center operations. In this paper, we present a framework of power budgeting, considering both computing power and cooling power, in data centers to maximize the system normalized performance (SNP) of the entire center under a total power budget. Maximizing the SNP for a given power budget is equivalent to maximizing the energy efficiency. We propose a method to partition the total power budget among the cooling and computing infrastructure in a self-consistent way, where the cooling power is sufficient to extract the heat of the computing power. Intertwinedly, we devise an optimal computing power budgeting technique based on dynamic programming algorithm to determine the optimal power caps for the individual servers such that the available power could be efficiently translated to performance improvements. The optimal computing budgeting technique leverages a proposed online throughput predictor based on performance counter measurements to estimate the change in throughput of heterogeneous workloads as a function of allocated server power caps. We demonstrate that our proposed power budgeting method outperforms previous methods by 3-4 percent in terms of SNP using our data center simulation environment. While maintaining the improvement of SNP, our method improve fairness at best by 57 percent. We also evaluate the performance of our method in power saving scenario and dynamic power budgeting case. Xin Zhan, Sherief Reda |
IEEE Trans. Computers | 2 |
| 2014 | ABACUS: A technique for automated behavioral synthesis of approximate computing circuitsabstractMany classes of applications, especially in the domains of signal and image processing, computer graphics, computer vision, and machine learning, are inherently tolerant to inaccuracies in their underlying computations. This tolerance can be exploited to design approximate circuits that perform within acceptable accuracies but have much lower power consumption and smaller area footprints (and often better run times) than their exact counterparts. In this paper, we propose a new class of automated synthesis methods for generating approximate circuits directly from behavioral-level descriptions. In contrast to previous methods that operate at the Boolean level or use custom modifications, our automated behavioral synthesis method enables a wider range of possible approximations and can operate on arbitrary designs. Our method first creates an abstract synthesis tree (AST) from the input behavioral description, and then applies variant operators to the AST using an iterative stochastic greedy approach to identify the optimal inexact designs in an efficient way. Our method is able to identify the optimal designs that represent the Pareto frontier trade-off between accuracy and power consumption. Our methodology is developed into a tool we call ABACUS, which we integrate with a standard ASIC experimental flow based on industrial tools. We validate our methods on three realistic Verilog-based benchmarks from three different domains - signal processing, computer vision and machine learning. Our tool automatically discovers optimal designs, providing area and power savings of up to 50% while maintaining good accuracy. Kumud Nepal, R. Iris Bahar, Sherief Reda |
DATE | 4 |
| 2014 | Thermal-aware layout planning for heterogeneous datacentersabstractCooling power represents a significant portion of total power consumption in datacenters. Heterogeneous datacenters deploy clusters of servers with different hardware configurations, each offering its own performance and power characteristics. We observe that heterogeneous datacenters offer a unique opportunity to reduce cooling power through appropriate planning. In this paper we formulate the problem of rack layout for planning of heterogeneous datacenters, where the goal is to identify the best locations of the server racks with different hardware capabilities to improve the supply temperatures of the CRAC units and the total cooling power. We provide optimal solutions that take into account the impact of varying utilizations of datacenters and job scheduling methods. Using state-of-the-art thermal modeling tools, we prove that our methods lead to datacenter layouts with significant improvements in cooling power reduction, between 15.5%-38.5% based on the datacenter utilizations and an average of 28.3% without any negative side effects. Xin Zhan, Sherief Reda |
ISLPED | 3 |
| 2014 | Novel Techniques for High-Sensitivity Hardware Trojan Detection Using Thermal and Power MapsabstractHardware Trojans are malicious alterations or injections of unwanted circuitry to integrated circuits (ICs) by untrustworthy factories. They render great threat to the security of modern ICs by various unwanted activities such as bypassing or disabling the security fence of a system, leaking confidential information, deranging, or destroying the entire chip. Traditional testing strategies are becoming ineffective since these techniques suffer from decreased sensitivity toward small Trojans because of oversized chip and large amount of process variation present in nanometer technologies. The production volume along with decreased controllability and observability to complex ICs internals make it difficult to efficiently perform Trojan detection using typical structural tests like path latency and leakage power. In this paper, we propose a completely new post-silicon multimodal approach using runtime thermal and power maps for Trojan detection and localization. Utilizing the novel framework, we propose two different Trojan detection methods involving 2-D principal component analysis. First, supervised thresholding in case training data set is available and second, unsupervised clustering which require no prior characterization data of the chip. We introduce 11 regularization in the thermal to power inversion procedure which improves Trojan detection accuracy. To characterize ICs accurately, we perform our experiments in presence of realistic CMOS process variation. Our experimental evaluations reveal that our proposed methodology can detect very small Trojans with 3-4 orders of magnitude smaller power consumptions than the total power usage of the chip, while it scales very well because of the spatial view to ICs internals by the thermal mapping. Abdullah Nazma Nowroz, Kangqiao Hu, Farinaz Koushanfar, Sherief Reda |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2014 | Fast Design Exploration for Performance, Power and Accuracy Tradeoffs in FPGA-Based AcceleratorsabstractThe ease-of-use and reconfigurability of FPGAs makes them an attractive platform for accelerating algorithms. However, accelerating becomes a challenging task as the large number of possible design parameters lead to different accelerator variants. In this article, we propose techniques for fast design exploration and multi-objective optimization to quickly identify both algorithmic and hardware parameters that optimize these accelerators. This information is used to run regression analysis and train mathematical models within a nonlinear optimization framework to identify the optimal algorithm and design parameters under various objectives and constraints. To automate and improve the model generation process, we propose the use of L 1 -regularized least squares regression techniques.We implement two real-time image processing accelerators as test cases: one for image deblurring and one for block matching. For these designs, we demonstrate that by sampling only a small fraction of the design space (0.42% and 1.1%), our modeling techniques are accurate within 2%--4% for area and throughput, 8%--9% for power, and 5%--6% for arithmetic accuracy. We show speedups of 340× and 90× in time for the test cases compared to brute-force enumeration. We also identify the optimal set of parameters for a number of scenarios (e.g., minimizing power under arithmetic inaccuracy bounds). Onur Ulusel, Kumud Nepal, R. Iris Bahar, Sherief Reda |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2013 | High-throughput TSV testing and characterization for 3D integration using thermal mappingabstractWe propose a new framework to detect structural defects and characterize the variability in the electrical resistance of through-silicon vias (TSVs) in 3D ICs. Our method offers a number of advantages that have been hard to achieve in the past. In particular, the proposed framework provides high throughput TSV testing at pre-bonding stage. A resistive liquid electrode is placed at the back side of the device to conduct electric current from TSVs. The current passing through TSVs leads to heat generation which can be captured by a remote, high-sensitivity thermal camera. The captured thermal signatures from the TSVs are then contrasted against reference thermal maps generated from known good die and/or electro-thermal simulations of models of good TSVs. A proposed automatic classification technique is capable of determining the status of TSVs based on their thermal signatures. We demonstrate the viability of the proposed technique using extensive simulation results on realistic TSV configurations. Kapil Dev, Gary L. Woods, Sherief Reda |
DAC | 3 |
| 2013 | Techniques for energy-efficient power budgeting in data centersabstractWe propose techniques for power budgeting in data centers, where a large power budget is allocated among the servers and the cooling units such that the aggregate performance of the entire center is maximized. Maximizing the performance for a given power budget automatically maximizes the energy efficiency. We first propose a method to partition the total power budget among the cooling and computing units in a self-consistent way, where the cooling power is sufficient to extract the heat of the computing power. Given the computing power budget, we devise an optimal computing budgeting technique based on knapsack-solving algorithms to determine the power caps for the individual servers. The optimal computing budgeting technique leverages a proposed on-line throughput predictor based on performance counter measurements to estimate the change in throughput of heterogeneous workloads as a function of allocated server power caps. We set up a simulation environment for a data center, where we simulate the air flow and heat transfer within the center using computational fluid dynamic simulations to derive accurate cooling estimates. The power estimates for the servers are derived from measurements on a real server executing heterogeneous workload sets. Our budgeting method delivers good improvements over previous power budgeting techniques. Xin Zhan, Sherief Reda |
DAC | 2 |
| 2013 | High-sensitivity hardware trojan detection using multimodal characterizationabstractVulnerability of modern integrated circuits (ICs) to hardware Trojans has been increasing considerably due to the globalization of semiconductor design and fabrication processes. The large number of parts and decreased controllability and observability to complex ICs internals make it difficult to efficiently perform Trojan detection using typical structural tests like path latency and leakage power. In this paper, we present new accurate methods for Trojan detection that are based upon post-silicon multimodal thermal and power characterization techniques. Our approach first estimates the detailed post-silicon spatial power consumption using thermal maps of the IC, then applies 2DPCA to extract features of the spatial power consumption, and finally uses statistical tests against the features of authentic ICs to detect the Trojan. To characterize real-world ICs accurately, we perform our experiments in presence of 20% – 40% CMOS process variation. Our results reveal that our new methodology can detect Trojans with 3–4 orders of magnitude smaller power consumptions than the total power usage of the chip, while it scales very well because of the spatial view to the ICs internals by the thermal mapping. Kangqiao Hu, Abdullah Nazma Nowroz, Sherief Reda, Farinaz Koushanfar |
DATE | 3 |
| 2013 | Mitigating dark-silicon problems using superlattice-based thermoelectric coolersabstractDark silicon is an emerging problem in multi-core processors, where it is not possible to enable all cores simultaneously because of either insufficient parallelism in software applications or because of high-spatial power densities that generate hot-spot constraints. Superlattice-based thermoelectric cooling (TEC) is a promising technology that offers large heat pumping capability and the ability to target hot spots of each core independently. In this paper, we devise novel system-level methods that address the two main sources of dark silicon using superlattice TECs. Our methods leverage the TECs in conjunction with dynamic voltage and frequency scaling and number of threads to maximize the performance of multi-core processor under thermal and power constraints. Using an experimental setup based on a quad-core processor, we provide an evaluation of the trade-offs among performance, temperature and power consumption arising from the use of superlattice-based TECs. Our results demonstrate the potential of this emerging cooling technology in mitigating dark silicon problems and in improving the performance of multi-core processors. Francesco Paterna, Sherief Reda |
DATE | 2 |
| 2013 | Power mapping and modeling of multi-core processorsabstractWe propose new techniques for post-silicon power mapping and modeling of multi-core processors using infrared imaging and performance counter measurements. An accurate finite-element modeling framework is used to capture the relationship between temperature and power, while compensating for the artifacts introduced from substituting traditional heat removal mechanisms with oil-based infrared-transparent cooling mechanisms. We use thermal conditioning techniques to build leakage power models for the die. Utilizing the power maps identified from infrared mapping, we develop empirical power models for different processor blocks based on the measurements from the performance monitoring counters (PMCs), and utilize the PMC-based models to analyze the transient power consumption. In our experiments, we capture thermal images from a quad-core processor under different workload conditions, and then we reconstruct the dynamic and leakage power maps for different blocks. Our results show good accuracy in mapping and modeling, revealing good insights into the trends of power consumption in multi-core processors. Kapil Dev, Abdullah Nazma Nowroz, Sherief Reda |
ISLPED | 3 |
| 2013 | vCap: Adaptive power capping for virtualized serversabstractPower capping on server nodes has become an essential feature in data centers for controlling energy costs and peak power consumption. More than half of the server nodes are virtualized in today's data centers; thus, providing a practical power capping technique for consolidated virtual environments is a significant research problem. This paper proposes a power capping technique, vCap, which makes resource allocation decisions to maximize the Quality-of-Service (QoS) while meeting the power constraints in virtualized servers that run multi-threaded applications. For a given set of applications, vCap first decides which applications to co-schedule based on application scalability and then optimizes the QoS in an application-aware manner for each VM by adaptively adjusting the CPU resources. Experiments on real-life multi-core servers show that vCap provides 12% higher energy efficiency in comparison to the state-of-the-art power capping techniques, while adhering to the power cap 92% of the time within a 2W error margin. Can Hankendi, Sherief Reda, Ayse K. Coskun |
ISLPED | 2 |
| 2013 | Post-silicon power mapping techniques for integrated circuits
Sherief Reda, Abdullah Nazma Nowroz, Ryan Cochran, Stefan Angelevski |
Integr. | 1 |
| 2013 | Power Mapping of Integrated Circuits Using AC-Based ThermographyabstractPost-silicon power validation is an important step in integrated circuit design and fabrication flow. It involves runtime power characterization of a fabricated chip under realistic loadings. The most versatile procedure for post-silicon power characterization involves capturing the thermal emissions from the back of the die and inverting the captured images to get power estimates. This process faces two major challenges: the spatial heat diffusion effect, which blurs the underlying power map, and measurement noise in the thermal imaging system. In this paper, we propose to use ac-based thermography, where ac excitation signals are applied to the chip instead of dc excitation signals, to improve post-silicon power mapping. We show that using ac excitation reduces the impact of flicker noise and spatial heat diffusion, which translates to significant improvements in power mapping accuracy. We perform a number of experiments using a test chip that can be programmed to control spatial and temporal power consumption. We use the test chip to analyze the noise in our thermal imaging system, and to quantify the improvements in power mapping attained from the proposed ac-based methodology. We elucidate the impact of the ac excitation frequency on both the signal-to-noise ratio and power mapping accuracy. We also demonstrate the basic applicability of our technique on a dual-core processor. Abdullah Nazma Nowroz, Gary L. Woods, Sherief Reda |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2012 | Fast Multi-Objective Algorithmic Design Co-Exploration for FPGA-based AcceleratorsabstractThe reconfigurability of Field Programmable Gate Arrays (FPGAs) makes them an attractive platform for accelerating algorithms. Accelerating a particular algorithm is a challenging task as the large number of possible algorithmic and hardware design parameters lead to different accelerator variant implementations, each with its own metrics such as performance, area, power, and arithmetic accuracy characteristics. To identify these parameters that optimize the accelerator for certain metrics, we propose techniques for fast design space exploration and non-linear multi-objective optimization (e.g., minimize power under arithmetic inaccuracy bounds). Our methodology samples a small part of the design space and uses measurements from the sampled implementations to train mathematical models for the different metrics. To automate and improve the model generation process, we propose the use of L1-regularized least squares regression techniques. To demonstrate the effectiveness of our approach, we implement a high-throughput real-time accelerator for image debluring. We demonstrate the accuracy (e.g., within 8% for power modeling) of our modeling techniques and their ability to identify the optimal accelerator designs with large speed-ups (340×) in comparison to brute-force enumeration. Kumud Nepal, Onur Ulusel, R. Iris Bahar, Sherief Reda |
FCCM | 4 |
| 2012 | Thermal prediction and adaptive control through workload phase detectionabstractElevated die temperature is a true limiter to the scalability of modern processors. With continued technology scaling in order to meet ever-increasing performance demands, it is no longer cost effective to design cooling systems that handle the worst-case thermal behaviors. Instead, cooling systems are designed to handle typical chip operation, while processors must detect and handle rare thermal emergencies. Most processors rely on measurements from integrated thermal sensors and dynamic thermal management (DTM) techniques in order to manage the trade-off between performance and thermal risk. Optimal management requires advanced knowledge of the thermal trajectory based on the current workload behaviors and operating conditions. In this work, we devise novel workload phase classification strategies that automatically discriminate among workload behaviors with respect to the thermal control response. We incorporate workload phase-detection and thermal models into a dynamic voltage and frequency scaling (DVFS) technique that can optimally control temperature during runtime based on thermal predictions. We demonstrate the effectiveness of our proposed techniques in predicting and adaptively controlling the thermal behavior of a real quad-core processor in response to a wide range of workloads. In comparison with state-of-the-art model predictive control (MPC) techniques in previous works on thermal prediction, we demonstrate a 5.8% improvement in instruction throughput with the same number of thermal violations. In comparison with simple proportional-integral (PI) feedback control techniques, we improve instruction throughput by 3.9%, while significantly reducing the number of thermal violations. Ryan Cochran, Sherief Reda |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2011 | Improved post-silicon power modeling using AC lock-in techniquesabstractThe objective of power modeling is to estimate the power consumption of integrated circuits under different workloads and variabilities. Post-silicon power modeling is an essential step for design validation and for building trustable pre-silicon power models and analyses. One popular approach for devising post-silicon power estimates is to translate the thermal emissions from the backside of the die into power estimates. Such approach faces a major physical challenge arising from spatial heat diffusion which blurs the resultant thermal images. The objective of this paper is to improve post-silicon power mapping by utilizing lock-in thermography techniques where AC signals, rather than DC signals, are used to excite the circuit blocks. We prove and demonstrate that using AC excitation sources reduces the extent of spatial heat diffusion. We devise a lock-in based thermal to power inversion methodology that maps spatial power consumption on a real chip. Using a custom test chip, we are to able to scientifically quantify and validate the improvements in power mapping attained from the proposed techniques. We show that our technique reduces the power mapping errors by more than half. Abdullah Nazma Nowroz, Gary L. Woods, Sherief Reda |
DAC | 3 |
| 2011 | Thermal and power characterization of field-programmable gate arraysabstractIn this paper we propose new techniques for thermal and power characterization of Field Programmable Gate Arrays (FPGAs) using infrared imaging techniques. For thermal characterization, we capture the thermal emissions from the backside of an FPGA chip during operation. We analyze the captured emissions and quantify the extent of thermal gradients and hot spots in FPGAs. Given that FPGAs are fabricated with no knowledge of the potential field designs, we propose soft sensing techniques that can combine the measurements of hard sensors to accurately estimate the temperatures where no sensors are embedded. For power characterization, we propose algorithmic techniques to invert the thermal emissions from FPGAs into spatial power estimates. We demonstrate how this technique can be used to produce spatial power maps of soft processors during operation. Abdullah Nazma Nowroz, Sherief Reda |
FPGA | 2 |
| 2011 | Identifying the optimal energy-efficient operating points of parallel workloadsabstractAs the number of cores per processor grows, there is a strong incentive to develop parallel workloads to take advantage of the hardware parallelism. In comparison to single-threaded applications, parallel workloads are more complex to characterize due to thread interactions and resource stalls. This paper presents an accurate and scalable method for determining the optimal system operating points (i.e., number of threads and DVFS settings) at runtime for parallel workloads under a set of objective functions and constraints that optimize for energy efficiency in multi-core processors. Using an extensive training data set gathered for a wide range of parallel workloads on a commercial multi-core system, we construct multinomial logistic regression (MLR) models that estimate the optimal system settings as a function of workload characteristics. We use L1-regularization to automatically determine the relevant workload metrics for energy optimization. At runtime, our technique determines the optimal number of threads and the DVFS setting with negligible overhead. Our experiments demonstrate that our method outperforms prior techniques with up to 51% improved decision accuracy. This translates to up to 10.6% average improvement in energy-performance operation, with a maximum improvement of 30.9%. Our technique also demonstrates superior scalability as the number of potential system operating points increases. Ryan Cochran, Can Hankendi, Ayse K. Coskun, Sherief Reda |
ICCAD | 4 |
| 2011 | Pack & Cap: adaptive DVFS and thread packing under power capsabstractThe ability to cap peak power consumption is a desirable feature in modern data centers for energy budgeting, cost management, and efficient power delivery. Dynamic voltage and frequency scaling (DVFS) is a traditional control knob in the tradeoff between server power and performance. Multi-core processors and the parallel applications that take advantage of them introduce new possibilities for control, wherein workload threads are packed onto a variable number of cores and idle cores enter low-power sleep states. This paper proposes Pack & Cap, a control technique designed to make optimal DVFS and thread packing control decisions in order to maximize performance within a power budget. In order to capture the workload dependence of the performance-power Pareto frontier, a multinomial logistic regression (MLR) classifier is built using a large volume of performance counter, temperature, and power characterization data. When queried during runtime, the classifier is capable of accurately selecting the optimal operating point. We implement and validate this method on a real quad-core system running the PARSEC parallel benchmark suite. When varying the power budget during runtime, Pack & Cap meets power constraints 82% of the time even in the absence of a power measuring device. The addition of thread packing to DVFS as a control knob increases the range of feasible power constraints by an average of 21% when compared to DVFS alone and reduces workload energy consumption by an average of 51.6% compared to existing control techniques that achieve the same power range. Ryan Cochran, Can Hankendi, Ayse K. Coskun, Sherief Reda |
MICRO | 4 |
| 2011 | Improved Thermal Tracking for Processors Using Hard and Soft Sensor Allocation TechniquesabstractHot spots are a major concern in high-end processors as they constrain performance and limit the lifetime of semiconductor chips. Using embedded thermal sensors, dynamic thermal management systems track the hot spots during runtime and adjust the performance and the cooling system of the processor when necessary. In many-core processors, the locations of hot spots vary spatially and temporally depending on the configuration of active cores and the workloads running on the cores. Our work includes both theoretical advances in sensor allocation techniques and experimental advances for thermal imaging of real processors. We propose a hard sensor allocation algorithm to determine the sensor locations where hot spots can be tracked accurately given a budget number of sensors. We also propose soft sensor computation techniques to alleviate design constraints on sensor locations and to further improve the resolution of hot spot tracking. The proposed soft sensing technique combines the measurements of the hard sensors in an optimal way to estimate the temperature at any desired location. We use infrared imaging methods to characterize the thermal behavior of a real dual-core processor during operation. We execute large number of workload configurations on the processor and track the locations and temperatures of hot spots during runtime. The thermal characterization data are then used as the input to our sensor allocation techniques. We demonstrate that our sensor allocation techniques improve significantly upon the previous results in the literature and provide accurate tracking of hot spots. Sherief Reda, Ryan Cochran, Abdullah Nazma Nowroz |
IEEE Trans. Computers | 1 |
| 2010 | Consistent runtime thermal prediction and control through workload phase detectionabstractElevated temperatures impact the performance, power consumption, and reliability of processors, which rely on integrated thermal sensors to measure runtime thermal behavior. These thermal measurements are typically inputs to a dynamic thermal management system that controls the operating parameters of the processor and cooling system. The ability to predict future thermal behavior allows a thermal management system to optimize a processor's operation so as to prevent the on-set of high temperatures. In this paper we propose a new thermal prediction method that leads to consistent results between the thermal models used in prediction and observed thermal sensor measurements, and is capable of accurately predicting temperature behavior with heterogenous workload assignment on a multicore platform. We devise an off-line analysis algorithm that learns a set of thermal models as a function of operating frequency and globally defined workload phases. We incorporate these thermal models into a dynamic voltage and frequency scaling (DVFS) technique that limits the maximum temperature during runtime. We demonstrate the effectiveness of our proposed system in predicting the thermal behavior of a real quad-core processor in response to different workloads. In comparison to a reactive thermal management technique, our predictive method dramatically reduces the number of thermal violations, the magnitude of thermal cycles, and workload runtimes. Ryan Cochran, Sherief Reda |
DAC | 2 |
| 2010 | Thermal monitoring of real processors: techniques for sensor allocation and full characterizationabstractThe increased power densities of multi-core processors and the variations within and across workloads lead to runtime thermal hot spots locations of which change across time and space. Thermal hot spots increase leakage, deteriorate timing, and reduce the mean time to failure. To manage runtime thermal variations, circuit designers embed within-die thermal sensors that acquire temperatures at few selected locations. The acquired temperatures are then used to guide runtime thermal management techniques. The capabilities of these techniques are essentially bounded by the spatial thermal resolution of the sensor measurements. In this paper we characterize temperature signals of real processors and demonstrate that on-chip thermal gradients lead to sparse signals in the frequency domain. We exploit this observation to (1) devise thermal sensor allocation techniques, and (2) devise signal reconstruction techniques that fully characterize the thermal status of the processor using the limited number of measurements from the thermal sensors. To verify the accuracy of our methods, we compare our temperature characterization results against thermal measurements acquired from a state-of-the-art infrared camera that captures the mid-band infrared emissions from the back of the die of a 45 nm dual-core processor. Our results show that our techniques are capable of accurately characterizing the temperatures of real processors. Abdullah Nazma Nowroz, Ryan Cochran, Sherief Reda |
DAC | 3 |
| 2010 | Post-silicon power characterization using thermal infrared emissionsabstractDesign-time power analysis is one of the most critical tasks conducted by chip architects and circuit designers. While computer-aided power analysis tools can provide power consumption estimates for various circuit blocks, these estimates can substantially deviate from the actual power consumption of working silicon chips. We propose a novel methodology that provides accurate, detailed post-silicon spatial power estimates using the thermal infrared emissions from the backside of silicon die. We theoretically and empirically demonstrate the inherent difficulties in thermal to power inversion. These difficulties arise from measurement errors and from the inherent spatial low-pass filtering associated with heat diffusion. To address these difficulties we propose new techniques from regularization theory to invert temperature to power. Furthermore, we propose new techniques to compute the emissivities and conductances required for any infrared to power inversion method. To verify our results, a programmable circuit of micro heaters is implemented to create any desired power pattern. The thermal emissions of different known injected power patterns are captured using a state-of-the-art infrared camera, and then our characterization techniques are applied to invert the thermal emissions to power. The estimated power patterns are validated against the injected power patterns to demonstrate the accuracy of our methodology. Ryan Cochran, Abdullah Nazma Nowroz, Sherief Reda |
ISLPED | 3 |
| 2009 | Spectral techniques for high-resolution thermal characterization with limited sensor dataabstractElevated chip temperatures are true limiters to the scalability of computing systems. Excessive runtime thermal variations compromise the performance and reliability of integrated circuits. To address these thermal issues, state-of-the-art chips have integrated thermal sensors that monitor temperatures at a few selected die locations. These temperature measurements are then used by thermal management techniques to appropriately manage chip performance. Thermal sensors and their support circuitry incur design overheads, die area, and manufacturing costs. In this paper, we propose a new direction for full thermal characterization of integrated circuits based on spectral Fourier analysis techniques. Application of these techniques to temperature sensing is based on the observation that die temperature is simply a space-varying signal, and that space-varying signals are treated identically to time-varying signals in signal analysis. We utilize Nyquist-Shannon sampling theory to devise methods that can almost fully reconstruct the thermal status of an integrated circuit during runtime using a minimal number of thermal sensors. We propose methods that can handle uniform and non-uniform thermal sensor placements. We develop an extensive experimental setup and demonstrate the effectiveness of our methods by thermally characterizing a 16-core processor. Our method produces full thermal characterization with an average absolute error of 0.6% using a limited number of sensors. Ryan Cochran, Sherief Reda |
DAC | 2 |
| 2009 | Analyzing the impact of process variations on parametric measurements: Novel models and applicationsabstractIn this paper we propose a novel statistical framework to model the impact of process variations on semiconductor circuits through the use of process sensitive test structures. Based on multivariate statistical assumptions, we propose the use of the expectation-maximization algorithm to estimate any missing test measurements and to calculate accurately the statistical parameters of the underlying multivariate distribution. We also propose novel techniques to validate our statistical assumptions and to identify any outliers in the measurements. Using the proposed model, we analyze the impact of the systematic and random sources of process variations to reveal their spatial structures. We utilize the proposed model to develop a novel application that significantly reduces the volume, time, and costs of the parametric test measurements procedure without compromising its accuracy. We extensively verify our models and results on measurements collected from more than 300 wafers and over 25 thousand die fabricated at a state-of-the-art facility. We prove the accuracy of our proposed statistical model and demonstrate its applicability towards reducing the volume and time of parametric test measurements by about 2.5 - 6.1times at absolutely no impact to test quality. Sherief Reda, Sani R. Nassif |
DATE | 1 |
| 2009 | High-performance, cost-effective heterogeneous 3D FPGA architecturesabstractIn this paper, we propose novel architectural and design techniques for three-dimensional field-programmable gate arrays (3D FPGAs) with Through-Silicon Vias (TSVs). We develop a novel design partitioning methodology that maps the heterogeneous computational resources of an FPGA into a number of die such that the total die area is minimized and the FPGA performance is maximized. Minimizing the total die area leads to direct manufacturing cost savings which is an important incentive to bring 3D technology to the fab and onto the market. An estimation framework is developed to assess the impact of silicon area utilized by 3D interconnect resources while taking into account the large area occupied by TSVs which is crucial to total die area of 3D FPGA. And in order to improve area and performance of 3D FPGA, we design a novel 3D switch box with bypass TSVs. We also analyze the impact of different partitioning strategies on die area and find the optimal number of die that gives the largest reductions in total die area while maximizing the performance. Using a well-developed simulation infrastructure, we show that our methodologies can achieve an average reduction of 27.7% in total die area with a reduced interconnect path delay of about 58%. Roto Le, Sherief Reda, R. Iris Bahar |
FPGA | 2 |
| 2009 | Central vs. distributed dynamic thermal management for multi-core processors: which one is better?abstractIn this paper we investigate and contrast two techniques to maximize the performance of multi-core processors under thermal constraints. The first technique is a distributed dynamic thermal management system that maximizes the total performance without exceeding given thermal constraints. In our scheme, each core adjusts its operating parameters, i.e., frequency and voltage, according to its temperature which is measured using integrated thermal sensors. We propose a novel controller that dynamically adapts the system to simultaneously avoid timing errors and thermal violations. For comparison purposes, we implement a second technique based on a runtime centralized, optimal system that uses combinatorial optimization techniques to calculate the optimal frequencies and voltages for the different cores to maximize the total throughput under thermal constraints. To empirically validate our techniques, we put together an extensive tool chain that incorporates thermal and power consumption simulators to characterize the performance of multi-core processors for a number of configurations ranging from 2 cores at 90 nm to 16 cores at 32 nm. Our results show that both investigated techniques are capable of delivering significant improvements (about 40% for 16 cores) over standard frequency and voltage planning techniques. From the results, we outline the main advantages and disadvantages of both techniques. Michael Kadin, Sherief Reda, Augustus K. Uht |
ACM Great Lakes Symposium on VLSI | 2 |
| 2009 | High-performance, cost-effective heterogeneous 3D FPGA architecturesabstractIn this paper, we propose novel architectural and design techniques for three-dimensional field-programmable gate arrays (3D FPGAs) with Through-Silicon Vias (TSVs). We develop a novel design partitioning methodology that maps the heterogeneous computational resources of an FPGA into a number of die such that the total die area is minimized and the FPGA performance is maximized. Minimizing the total die area leads to direct manufacturing cost savings which is an important incentive to bring 3D technology to the fab and onto the market. An estimation framework is developed to assess the impact of silicon area utilized by 3D interconnect resources while taking into account the large area occupied by TSVs which is crucial to total die area of 3D FPGAs. In order to improve area and performance of 3D FPGAs, we design a novel 3D switch box with bypass TSVs. We also analyze the impact of different partitioning strategies on die area and find the optimal number of die that gives the largest reductions in total die area while maximizing the performance. Using a well-developed simulation infrastructure, we show that our methodologies can achieve an average reduction of 27.7% in total die area with a reduced interconnect path delay of about 58%. Roto Le, Sherief Reda, R. Iris Bahar |
ACM Great Lakes Symposium on VLSI | 2 |
| 2009 | Reducing the leakage and timing variability of 2D ICcs using 3D ICsabstractThis paper examines the ramifications of using 3D integration technology on the leakage and timing variability of integrated circuits. We develop models that estimate the outcome of mapping a 2D design onto a 3D stack from a process variation perspective. We statistically prove and experimentally demonstrate that 3D integration is a useful technique to combat process variations even if the die/wafers layers involved in 3D stacks are integrated blindly without any parametric tests prior to integration. We further show that if individual die parametric testing information is available, then it is possible to drastically reduce the impact of process variations. We develop fast, near optimal integration strategies based on recursive matching techniques. Our results show that 3D integration can reduce the variability in leakage and timing of planar ICs by around 50% without any testing and by more than 90% with additional test requirements. Sherief Reda, Aung Si, R. Iris Bahar |
ISLPED | 1 |
| 2009 | Maximizing the Functional Yield of Wafer-to-Wafer 3-D IntegrationabstractThree-dimensional integrated circuit technology with through-silicon vias offers many advantages, including improved form factor, increased circuit performance, robust heterogenous integration, and reduced costs. Wafer-to-wafer integration supports the highest possible density of through-silicon vias and highest throughput; however, in contrast to die-to-wafer integration, it does not benefit from the ability to bond only tested and diced good die. In wafer-to-wafer integration, wafers are entirely bonded together, which can unintentionally integrate a bad die from one wafer to a good die from another wafer reducing the yield. In this paper, we propose solutions that maximize the yield of wafer-to-wafer 3-D integration, assuming that the individual die can be tested on the wafers before bonding. We exploit some of the available flexibility in the integration process, and propose wafer assignment algorithms that maximize the number of good 3-D ICs. Our algorithms range from scalable, fast heuristics to optimal methods that exactly maximize the yield of wafer-to-wafer 3-D integration. Using realistic defect models and yield simulations, we demonstrate the effectiveness of our methods up to large numbers of wafer stacks. Our results demonstrate that it is possible to significantly improve the yield in comparison to yield-oblivious wafer assignment methods. Sherief Reda, Gregory Smith 0002, Larry Smith 0004 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2008 | Within-die process variations: How accurately can they be statistically modeled?abstractWithin-die process variations arise during integrated circuit (IC) fabrication in the sub-100nm regime. These variations are of paramount concern as they deviate the performance of ICs from their designers' original intent. These deviations reduce the parametric yield and revenues from integrated circuit fabrication. In this paper we provide a complete treatment to the subject of within-die variations. We propose a scan-chain based system, vMeter, to extract within-die variations in an automated fashion. We implement our system in a sample of 90 nm chips, and collect the within-die variations data. Then we propose a number of novel statistical analysis techniques that accurately model the within-die variation trends and capture the spatial correlations. We propose the use of maximum-likelihood techniques to find the required parameters to fit the model to the data. The accuracy of our models is statistically verified through residual analysis and variograms. Using our successful modeling technique, we propose a procedure to generate synthetic within-die variation patterns that mimic, or imitate, real silicon data. Brendan Hargreaves, Henrik Hult, Sherief Reda |
ASP-DAC | 3 |
| 2008 | Frequency and voltage planning for multi-core processors under thermal constraintsabstractClock frequency and transistor density increases have resulted in elevated chip temperatures. In order to meet temperature constraints while still exploiting the performance opportunities enabled by continued scaling, chip designers have migrated towards multi-core architectures. Multi-core architectures use multiple cores running at moderate clock frequencies to run several threads concurrently, which increases overall system throughput. In this work, we propose novel methods to find the optimal operating parameters, i.e., frequency and voltage, that maximize a multi-core system throughput under thermal constraints. By adjusting core clock frequencies and voltages, on-chip power dissipation can be spatially and temporally distributed to maximize the chippsilas physical performance during runtime. We propose a simple, yet efficient model that accurately characterize the effects that changes in clock frequency and voltage have on on-chip temperatures. Using the model, we find the optimal operating conditions for the following scenarios: (1) standard processor performance, where various cores operate using identical operating parameters, (2) optimal processor performance where each core can have its own frequency and voltage, and (3) optimal processor performance with thread priorities, where each core runs a thread of varied importance. We run several experiments across six different technology nodes to validate the work, assuring that our models and methods are accurate. Our methods demonstrate the total physical performance of a multi-core system can be increased by up to 33.4% without violating the maximum temperature constraints. Michael Kadin, Sherief Reda |
ICCD | 2 |
| 2008 | Frequency planning for multi-core processors under thermal constraintsabstractThe objectives of this paper are (1) to develop a frequency planning methodology that maximizes the total performance of multi-core processors and that limits their maximum temperature as specified by the design constraints; and (2) to establish the implications of technology scaling on the performance limits of multi-core processors. Given the intricate designs and workloads of multi or many-core processors, it is computationally exhaustive to develop models that accurately calculate the temperature and performance of a given processor under various operating conditions. To abstract the underlying design complexity, we propose the use of supervised machine learning techniques to develop versatile models that capture the thermal characterization of multi-core processors under various input conditions and workloads. We then use the developed models to create a framework where various design constraints and objectives are expressed and solved using combinatorial optimization techniques. Using established power modeling and thermal simulation tools, we show that it is possible to boost the performance of multi-core processors by up to 11.4% at no impact to the maximum temperature. Michael Kadin, Sherief Reda |
ISLPED | 2 |
| 2008 | Parametric yield management for 3D ICs: Models and strategies for improvementabstractThree-Dimensional (3D) Integrated Circuits (ICs) that integrate die with Through-Silicon Vias (TSVs) promise to continue system and functionality scaling beyond the traditional geometric 2D device scaling. 3D integration also improves the performance of ICs by reducing the communication time between different chip components through the use of short TSV-based vertical wires. This reduction is particularly attractive in processors where it is desirable to reduce the access time between the main logic die and the L2 cache or the main memory die. Process variations in 2D ICs lead to a drop in parametric yield (as measured by speed, leakage and sales profits), which forces manufacturers to speed bin their chips and to sell slow chips at reduced prices. In this paper we develop a model to quantify the impact of process variations on the parametric yield of 3D ICs, and then we propose a number of integration strategies that use a graph-theoretic framework to maximize the performance, parametric yield and profits of 3D ICs. Comparing our proposed strategies to current yield-oblivious methods, it is demonstrated that it is possible to increase the number of 3D ICs in the fastest speed bins by almost 2×, while simultaneously reducing the number of slow ICs by 29.4%. This leads to an improvement in performance by up to 6.45% and an increase of about 12.48% in total sales revenue using up-to-date market price models. Cesare Ferri, Sherief Reda, R. Iris Bahar |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2007 | Strategies for improving the parametric yield and profits of 3D ICsabstractThree-Dimensional (3D) Integrated Circuits (ICs) that integrate die with Through-Silicon Vias (TSVs) promise to continue system and functionality scaling beyond the traditional geometric 2D device scaling. 3D integration also improves the performance of ICs by reducing the communication time between different chip components through the use of short TSV-based vertical wires. This reduction is particularly attractive in processors where it is desirable to reduce the access time between the main logic die and the L2 cache or the main memory die. Process variations in 2D ICs lead to a drop in parametric yield (as measured by speed, leakage and sales profits), which forces manufacturers to speed bin their chips and to sell slow chips at reduced prices. In this paper we develop a model to quantify the impact of process variations on the parametric yield of 3D ICs, and then we propose a number of integration strategies that use a graph-theoretic framework to maximize the performance, parametric yield and profits of 3D ICs. Comparing our proposed strategies to current yield-oblivious methods, it is demonstrated that it is possible to increase the number of 3D ICs in the fastest speed bins by almost 2×, while simultaneously reducing the number of slow ICs by 29.4%. This leads to an improvement in performance by up to 6.45% and an increase of about 12.48% in total sales revenue using up-to-date market price models. Cesare Ferri, Sherief Reda, R. Iris Bahar |
ICCAD | 2 |
| 2007 | Hardware libraries: An architecture for economic acceleration in soft multi-core environmentsabstractIn single processor architectures, computationally- intensive functions are typically accelerated using hardware accelerators, which exploit the concurrency in the function code to achieve a significant speedup over software. The increased design constraints from power density and signal delay have shifted processor architectures in general towards multi-core designs. The migration to multi-core designs introduces the possibility of sharing hardware accelerators between cores. In this paper, we propose the concept of a hardware library, which is a pool of accelerated functions that are accessible by multiple cores. We find that sharing provides significant reductions in the area, logic usage and leakage power required for hardware acceleration. Contention for these units may exist in certain cases; however, the savings in terms of chip area are more appealing to many applications, particularly the embedded domain. We study the performance implications for our proposal using various multi-core arrangements, with actual implementations in FPGA fabrics. FPGAs are particularly appealing due to their cost effectiveness and the attained area savings enable designers to easily add functionality without significant chip revision. Our results show that is possible to save up to 37% of a chip's available logic and interconnect resources at a negligible impact (< 3%) to the performance. David Meisner, Sherief Reda |
ICCD | 2 |
| 2006 | Effective linear programming based placement methodsabstractLinear programming (LP) based methods are attractive for solving the placement problem because of their ability to model Half-Perimeter Wirelength (HPWL) and timing. However, it has been technically difficult to model overlaps in LP. This difficulty in modeling overlaps restricted the domain of LP-based methods to incremental placers, where LP is used to calculate the optimal locations of a small subset of cells with no regard to overlaps. In this paper, we enlarge the scope of LP-based methods from just operating on a small subset of cells to operating on all cells of a functional block circuit. We show how to model, reduce and prevent overlaps in LP-based placement flows. We use our ideas to construct (1) a global optimal whitespace allocator, and (2) a global overlap remover and cell spreader. We also modify our methods to fit in a timing-driven placement flow. Compared to our default industrial flow, our results show an improvement by an average of 7.64% in wirelength, and by an average of 21% in total negative slack. Furthermore, we conduct a benchmarking study, where we surprisingly show that academic placers fail to consistently produce good results on relatively small functional blocks. Sherief Reda, Amit Chowdhary |
ISPD | 1 |
| 2006 | Computer-Aided Optimization of DNA Array Design and ManufacturingabstractDNA probe arrays, or DNA chips, have emerged as a core genomic technology that enables cost-effective gene expression monitoring, mutation detection, single nucleotide polymorphism analysis, and other genomic analyses. DNA chips are manufactured through a highly scalable process called very large-scale immobilized polymer synthesis (VLSIPS) that combines photolithographic technologies adapted from the semiconductor industry with combinatorial chemistry. Commercially available DNA chips contain more than half a million probes and are expected to exceed 100 million probes in the next generation. This paper is one of the first attempts to apply very large scale integration (VLSI) computer-aided design methods to the physical design of DNA chips, where the main objective is to minimize total border cost (i.e., the number of nucleotide mismatches between adjacent sites). By exploiting analogies between manufacturing processes for DNA arrays and for VLSI chips, the authors demonstrate the potential for transfer of methodologies from the 40-year-old field of electronic design automation to the newer DNA array design field. The main contributions of this paper are the following. First, it proposes several partitioning-based algorithms for DNA probe placement that improve solution quality by over 4% compared to best previously known methods. Second, it gives a new design flow for DNA arrays, which enhances current methodologies by adding flow awareness to each optimization step and introducing feedback loops. Third, it proposes solution methods for new formulations integrating multiple design steps, including probe selection, placement, and embedding. Finally, it introduces new techniques to experimentally evaluate the scalability and suboptimality of existing and newly proposed probe placement algorithms. Interestingly, the authors find that DNA placement algorithms appear to have better suboptimality properties than those recently reported for VLSI placement algorithms [C.C. Chang et al., Optimality and scalability study of existing placement algorithms, Proc. Asia South-Pacific Design Automation Conf., Kitakyushu, Japan, p.621-7, Jan. 2003; J. Cong et al., Optimality, scalability and stability study of partitioning and placement algorithms, Proc. Int. Symp. Physical Design (ISPD), Monterey, CA, p.88-94, 2003] Andrew B. Kahng, Ion I. Mandoiu, Sherief Reda, Xu Xu 0001, Alex Zelikovsky |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2006 | New and improved BIST diagnosis methods from combinatorial Group testing theoryabstractWe examine the general problem of built-in-self-test (BIST) diagnosis in digital logic systems. The BIST diagnosis problem has applications that include identification of erroneous test vectors, faulty scan cells, and faulty items. We develop an abstract model of this problem and show a fundamental correspondence to the well-established subject of combinatorial group testing (CGT) (D. Du and F. K. Hwang, Combinatorial Group Testing and Its Applications, 1994). We exploit this new perspective to 1) link existing BIST diagnosis techniques to CGT techniques and provide further insights into existing diagnosis algorithms, 2) improve the performance of diagnosis algorithms, and 3) develop new techniques to address the BIST diagnosis problem. Using the ISCAS'89 benchmarks, we empirically demonstrate the effectiveness of our proposed techniques over existing BIST diagnosis techniques. The vastness of the CGT literature suggests that further improvements from existing research in CGT may be obtained. Andrew B. Kahng, Sherief Reda |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2006 | Wirelength minimization for min-cut placements via placement feedbackabstractThe advent of strong multilevel partitioners has made top-down min-cut placers a favored choice for modern placer implementations. Terminal propagation is an important step in min-cut placers because it translates partitioning results into global-placement wirelength assumptions. In this work, the repartitioning problem is carefully reexamined (Proc. ACM/IEEE Int. Symp. Physical Design, p. 18, 1997) in the context of terminal propagation and studied in an in-depth manner. Abstractly, it was observed that in repartitioning, future cell locations are used for present terminal propagations and that this can be conceptually regarded as a form of placement feedback. This concept was utilized to achieve accurate terminal propagation via feedback iteration and controller insertion to fine-tune the feedback response. This yields substantial reductions in placement wirelength. Implementing our approach in Capo [version 8.7 (Proc. ACM/IEEE Design Automation Conf., p. 477, 2000 and GSRC Bookshelf)] and applying it to standard benchmark circuits yields up to 14% wirelength reductions for the IBM benchmarks with an average improvement of 5.5% and up to 10% reductions for the Peko benchmarks with an average improvement of 5.37%. Experiments also show consistent improvements for routed wirelength, yielding up to 9% wirelength reductions and 5.8% average reduction with acceptable increase in placement runtime. In practice, the method proposed significantly improves routability without building congestion maps and also reduces the number of vias Andrew B. Kahng, Sherief Reda |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2006 | Zero-Change Netlist Transformations: A New Technique for Placement BenchmarkingabstractIn this paper, the authors introduce the concept of zero-change netlist transformations (ZCNTs) to: 1) quantify the suboptimality of existing placers on artificially constructed instances and 2) "partially" quantify the suboptimality of placers on synthesized netlists from arbitrary netlists by giving lower bounds to the suboptimality gap. Given a netlist and its placement from a placer, a class of netlist transformations that synthesizes a different netlist from the given netlist is formally defined, but yet the new netlist has the same half-perimeter wire length (HPWL) on the given placement. Furthermore, and more importantly, the optimal HPWL value of the new netlist is no less than that of the original netlist. By applying the transformations and reexecuting the placer, any deviation in HPWL as a lower bound to the gap from the optimal HPWL value of the new synthesized netlist can be interpreted. The transformations allow us to: 1) increase the cardinality of hyperedges; 2) reduce the number of hyperedges; and 3) increase the number of two-pin edges, while maintaining the placement HPWL constant. It is developed here methods that apply ZCNTs to synthesize netlists having typical netlist statistics. Furthermore, an approach to estimate the suboptimality of other metrics, such as rectilinear minimum-spanning tree (RMST) and minimum-Steiner tree, is extended. Using these transformations, the suboptimality of some of the existing academic placers (FengShui, Capo, mPL, Dragon) is studied on synthesized netlists from the IBM benchmarks with instances ranging from 10k to 210k placeable instances. The results show that current placers exhibit suboptimal behavior to ZCNTs with varying degree according to the placer. Systematic suboptimality deviations in HPWL and RMST are displayed on the synthesized netlists from IBM (version 1) benchmarks. The specific nature of the transformations points out troublesome netlist structures and possible directions for improvement in the existing placers Andrew B. Kahng, Sherief Reda |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2006 | A Fast Hierarchical Quadratic Placement AlgorithmabstractPlacement is a critical component of today's physical-synthesis flow with tremendous impact on the final performance of very large scale integration (VLSI) designs. Unfortunately, it accounts for a significant portion of the overall physical-synthesis runtime. With the complexity and the netlist size of today's VLSI design growing rapidly, clustering for placement can provide an attractive solution to manage affordable placement runtimes. However, such clustering has to be carefully devised to avoid any adverse impact on the final placement solution quality. This paper presents how to apply clustering and unclustering strategies to an analytic top-down placer to achieve large speedups without sacrificing (and sometimes even enhancing) the solution quality. The authors' new bottom-up clustering technique, called the best choice (BC), operates directly on a circuit hypergraph and repeatedly clusters the globally best pair of objects. Clustering score manipulation using a priority-queue (PQ) data structure enables identification of the best pair of objects whenever clustering is performed. To improve the runtime of PQ-based BC clustering, the authors proposed a lazy-update technique for faster updates of the clustering score with almost no loss of the solution quality. A number of effective methods for clustering score calculation, balancing cluster sizes, handling of fixed blocks, and area-based unclustering strategy are discussed. The effectiveness of the resulting hierarchical analytic placement algorithm is tested on several large-scale industrial benchmarks with mixed-size fixed blocks. Experimental results are promising. Compared to the flat analytic placement runs, the hierarchical mode is 2.1 times faster, on the average, with a 1.4% wire-length improvement. Gi-Joon Nam, Sherief Reda, Charles J. Alpert, Paul G. Villarrubia, Andrew B. Kahng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2005 | Power-aware placementabstractLowering power is one of the greatest challenges facing the IC industry today. We present a power-aware placement method that simultaneously performs (1) activity-based register clustering that reduces clock power by placing registers in the same leaf cluster of the clock trees in a smaller area and (2) activity-based net weighting that reduces net switching power by assigning a combination of activity and timing weights to the nets with higher switching rates or more critical timing. The method applies to designs with multiple clocks and gated clocks. We implemented the method and obtained experimental results on 8 real-world designs after placement, routing, extraction and analysis. The power-aware placement method achieved on average 25.3% and 11.4% reduction in net switching power and total power respectively, with 2.0% timing, 1.2% cell area and 11.5% runtime impact. This method has been incorporated into a commercial physical design tool. Yongseok Cheon, Pei-Hsin Ho, Andrew B. Kahng, Sherief Reda, Qinke Wang |
DAC | 4 |
| 2005 | Intrinsic shortest path length: a new, accurate a priori wirelength estimatorabstractA priori wirelength estimation is concerned with predicting various wirelength characteristics before placement. In this work we propose a novel, accurate estimator of net lengths. We observe that in "good" placements, the length of a net is very strongly correlated with the numbers of nets in the shortest paths connecting node pairs of the net, when each shortest path is computed under the restriction that the net itself does not exist. We refer to this as the net's intrinsic shortest path length (ISPL). Using ISPL as a wirelength estimator has several advantages: (1) it transparently handles multi-pin nets and is a strong predictor of their length; (2) it strongly correlates with the average netlist wirelength; (3) it has a distribution that is similar to that of wirelength; and (4) it acts as a good predictor for individual net lengths. Based on ISPLs, we characterize VLSI netlists with a single value and develop an intuitive, empirical link between our proposed value and the Rent parameter. We also analytically model the relationship between ISPL and wirelength, and use ISPLs in two practical applications: a priori total wirelength estimation and a priori global interconnect prediction. Andrew B. Kahng, Sherief Reda |
ICCAD | 2 |
| 2005 | Architecture and details of a high quality, large-scale analytical placerabstractModern design requirements have brought additional complexities to netlists and layouts. Millions of components, whitespace resources, and fixed/movable blocks are just a few to mention in the list of complexities. With these complexities in mind, placers are faced with the burden of finding an arrangement of placeable objects under strict wirelength, timing, and power constraints. In this paper we describe the architecture and novel details of our high quality, large-scale analytical placer. The performance of our placer has been recently recognized in the recent ISPD-2005 placement contest, and in this paper we disclose many of the technical details that we believe are key factors to its performance. We describe (i) a new clustering architecture, (ii) a dynamically adaptive analytical solver, and (iii) better legalization schemes and novel detailed placement methods. We also provide extensive experimental results on a number of benchmark sets. On average, our results are better than the best published results by 3%, 14%, and 6% for the IBM ISPD '04, ICCAD '04, and ISPD '05 benchmark sets respectively. One of the goals of this paper is to also provide enough details to enable possible future replication of our methods. Andrew B. Kahng, Sherief Reda, Qinke Wang |
ICCAD | 2 |
| 2005 | A semi-persistent clustering technique for VLSI circuit placementabstractPlacement is a critical component of today's physical synthesis flow with tremendous impact on the final performance of VLSI designs. However, it accounts for a significant portion of the over-all physical synthesis runtime. With complexity and netlist size of today's VLSI design growing rapidly, clustering for placement can provide an attractive solution to manage affordable placement runtime. Such clustering, however, has to be carefully devised to avoid any adverse impact on the final placement solution quality. In this paper we present a new bottom-up clustering technique, called best-choice, targeted for large-scale placement problems. Our best-choice clustering technique operates directly on a circuit hypergraph and repeatedly clusters the globally best pair of objects. Clustering score manipulation using a priority-queue data structure enables us to identify the best pair of objects whenever clustering is performed. To improve the runtime of priority-queue-based best-choice clustering, we propose a lazy-update technique for faster updates of clustering score with almost no loss of solution quality. We also discuss a number of effective methods for clustering score calculation, balancing cluster sizes, and handling of fixed blocks. The effectiveness of our best-choice clustering methodology is demonstrated by extensive comparisons against other standard clustering techniques such as Edge-Coarsening [12] and First-Choice [13]. All clustering methods are implemented within an industrial placer CPLACE [1] and tested on several industrial benchmarks in a semi-persistent clustering context. Charles J. Alpert, Andrew B. Kahng, Gi-Joon Nam, Sherief Reda, Paul G. Villarrubia |
ISPD | 4 |
| 2005 | Evaluation of placer suboptimality via zero-change netlist transformationsabstractIn this paper we introduce the concept of zero-change transformations to quantify the suboptimality of existing placers. Given a netlist and its placement from a placer, we formally define a class of netlist transformations that produce different netlists from the given netlist but have the same Half-Perimeter Wire Length (HPWL). Furthermore, the optimal HPWL value of the new netlists is no less than that of the original netlist. By applying our transformations and re-executing the placer, we can interpret any deviation in HPWL as a lower bound to the deviation from the optimal HPWL value. Such deviation is a measure of suboptimality. Using these transformations, the suboptimality of several existing academic and industrial placers is studied on the IBM benchmarks. Our results show that current placers are sub-optimal for zero-change transformations with deviations in HPWL by up to 32% on the IBM (version 1) benchmarks. The specific nature of our transformations also pinpoints possible directions for improvement in existing placers. Andrew B. Kahng, Sherief Reda |
ISPD | 2 |
| 2005 | APlace: a general analytic placement frameworkabstractWe streamline and extend APlace, the general analytic placement engine based on ideas of Naylor et al. [7] and described in [3, 4, 5]. Previous work explored the adaptability of APlace to multiple contexts with good quality of results. For example, the framework was extended to traditional wirelength-driven standard-cell placement in [3, 5], achieving good results in placed HPWL and routed final wire-length. The framework was also extended to top-down multilevel placement, congestion-directed placement, mixed-size placement, timing-driven placement, I/O-core co-placement and constraint handling for mixed-signal contexts [3, 4, 5]. In this work, we have modified the implementation of APlace for speed and scalability. Improvements have been made in clustering, legalization and detailed placement strategies, as well as via a distributable solution framework for both global and detailed placement phases. Andrew B. Kahng, Sherief Reda, Qinke Wang |
ISPD | 2 |
| 2004 | Combinatorial group testing methods for the BIST diagnosis problem
Andrew B. Kahng, Sherief Reda |
ASP-DAC | 2 |
| 2004 | Placement feedback: a concept and method for better min-cut placementsabstractThe advent of strong multi-level partitioners has made topdown min-cut placers a favored choice for modern placer implementations. We examine terminal propagation, an important step in min-cut placers, because it is responsible for translating partitioning results into global placement wirelength assumptions. In this work, we identify a previously overlooked problem - ambiguous terminal propagation - and propose a solution based on the concept of feedback from automatic control systems. Implementing our approach in Capo (version 8.7 [5, 10]) and applying it to standard benchmark circuits yields up to 14% wirelength reductions for the IBM benchmarks and 10% reductions for PEKO instances. Experiments also show consistent improvements for routed wirelength, yielding up to 9% wirelength reductions with practical increase in placement runtime. In addition, our method significantly improves routability without building congestion maps, and reduces the number of vias. Andrew B. Kahng, Sherief Reda |
DAC | 2 |
| 2004 | Boosting: Min-Cut Placement with Improved Signal DelayabstractIn this work we improve top-down min-cut placers in the context of timing closure. Using the concept of boosting factors, we adjust net weights according to net spans, so as to reduce the quadratic wirelength. Our method is generic and does not involve any timing analysis during or prior to placement. In essence, we skew the netlength distribution produced by a min-cut placer so as to decrease the number of long nets, with minimal impact on the overall wirelength. Empirically this approach does not significantly affect runtime, but reduces the worst negative slack and total negative slack of industrial benchmarks by up to 70% compared to Capo and a leading industrial placer. Andrew B. Kahng, Igor L. Markov, Sherief Reda |
DATE | 3 |
| 2004 | On legalization of row-based placementsabstractCell overlaps in the results of global placement are guaranteed to prevent successful routing. However, common techniques for fixing these problems may endanger routing in a different way --- through increased wirelength and congestion. We evaluate several such techniques with routability of row-based placements in mind, and propose new ones that, in conjunction with our detail placer, improve overall routability and routed wirelength. Our generic two-phase approach for resolving illegal placements calls for (i) balancing the numbers of cells in rows, (ii) removing overlaps within rows through a generic dynamic programming procedure. Relevant objectives include minimum total perturbation, minimum wirelength increase and minimum maximum movement. Additionally, we trace cell overlaps in min-cut placement to vertical cuts and show that, if bisection cut directions are varied, overlaps anti-correlate with improved wirelength.Empirical validation is performed using placers Capo and Cadence QPlace, followed by various legalizers and detail placers, with subsequent routing by Cadence WarpRoute. We use a number of IBMv2 benchmarks with routing information. Our legalizer reduces both Capo and QPlace placements' wirelength by up to 4% compared to results of Capo legalized by Cadence's QPlace in the ECO mode. Andrew B. Kahng, Igor L. Markov, Sherief Reda |
ACM Great Lakes Symposium on VLSI | 3 |
| 2003 | Evaluation of Placement Techniques for DNA Probe Array Layout
Andrew B. Kahng, Ion I. Mandoiu, Sherief Reda, Xu Xu 0001, Alex Zelikovsky |
ICCAD | 3 |
| 2003 | Design Flow Enhancements for DNA ArraysabstractDNA probe arrays have recently emerged as one of the core genomic technologies. Exploiting analogies between manufacturing processes for DNA arrays and for VLSI chips, we demonstrate the potential for transfer of methodologies from the 40-year old field of electronic design automation to the newer DNA array design field. Our main contributions are the following. (1) We give a new design flow for DNA arrays which enhances current methodologies by adding flow-awareness to each optimization step and introducing feedback loops. (2) We propose solution methods for new formulations integrating multiple design steps, including probe selection, placement, and embedding. (3) We give results of a comprehensive experimental study showing that significant improvements in solution quality can be achieved by using the enhanced methodologies. Andrew B. Kahng, Ion I. Mandoiu, Sherief Reda, Xu Xu 0001, Alex Zelikovsky |
ICCD | 3 |
| 2003 | Engineering a scalable placement heuristic for DNA probe arraysabstractDesign of DNA arrays for very large-scale immobilized polymer synthesis (VLSIPS) [8] seeks to minimize effects of unintended illumination during mask exposure steps. [9, 14] formulate this requirement as the Border Minimization Problem and give methods for placement (at array sites) and embedding (in the mask sequence) of probes in both synchronous and asynchronous regimes. These previous methods do not address several practical details of the application and, more critically, are not scalable to the O(108) probes contemplated for next-generation probe arrays. In this work, we make two main contributions: Andrew B. Kahng, Ion I. Mandoiu, Pavel A. Pevzner, Sherief Reda, Alex Zelikovsky |
RECOMB | 4 |
| 2002 | Reducing Test Application Time Through Test Data Mutation EncodingabstractIn this paper we propose a new compression algorithm geared to reduce the time needed to test scan-based designs. Our scheme compresses the test vector set by encoding the bits that need to be flipped in the current test data slice in order to obtain the mutated subsequent test data slice. Exploitation of the overlap in the encoded data by effective traversal search algorithms results in drastic overall compression. The technique we propose can be utilized as not only a stand-alone technique but also can be utilized on test data already compressed, extracting even further compression. The performance of the algorithm is mathematically analyzed and its merits experimentally confirmed on the larger examples of the ISCAS '89 benchmark circuits. Sherief Reda, Alex Orailoglu |
DATE | 1 |
| 2002 | Border Length Minimization in DNA Array Design
Andrew B. Kahng, Ion I. Mandoiu, Pavel A. Pevzner, Sherief Reda, Alex Zelikovsky |
WABI | 4 |
| 2001 | Combinational equivalence checking using Boolean satisfiability and binary decision diagramsabstractMost recent combinational equivalence checking techniques are based on exploiting circuit similarity. In this paper, we focus on circuits with no internal equivalent nodes or after internal equivalent nodes have been identified and merged. We present a new technique integrating Boolean satisfiability and binary decision diagrams. The proposed approach is capable of solving verification instances that neither of the previous techniques was capable of solving. The efficiency of the proposed approach is shown through its application on hard to prove industrial circuits and the ISCAS'85 benchmark circuits. Sherief Reda, Ashraf Salem |
DATE | 1 |
| 2000 | M-CHECK: a multiple engine combinational equivalence checkerabstractSingle engine equivalence checkers have been successfully used in efficiently verifying large designs. However, they also showed inefficiency in proving specific types of circuits. New trends tend to combine multiple equivalence checking engines into a single framework, so that they can verify different kinds of designs more efficiently. In this paper, M-CHECK, a multiple engine equivalence checker, is presented. The proposed checker verifies the combinational circuits using three checking engines: Binary Decision Diagram (BDD) engine, Boolean Satisfiability (SAT) engine, and a mixed BDD SAT engine. Also, the tool is capable of iteratively reducing the design through identifying of equivalent node pairs. The results of our methodology are presented on the ISCAS-85 benchmark circuits. Sherief Reda, Ayman Wahba, Ashraf Salem |
ISCAS | 1 |