VLDB 2026 Research / reviewers in the wild / expert
Fadi J. Kurdahi
dblp:19/5115
· DBLP profile ↗
118ranked-venue papers
8as first author
17since 2021 · last 2026
0000-0002-6982-365XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 102 · 8 first-author · 13 since 2021Software engineering, systems software and programming languages · 13 · 4 since 2021Computer networks · 5Graphics, computer vision, multimedia, augmented reality and games · 3Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Kolmogorov-Arnold Networks for Autonomous Driving: A Hardware-in-the-Loop Comparison with Conventional Deep Neural NetworksabstractAutonomous driving systems require AI models that balance accuracy, efficiency, and interpretability for practical deployment. This paper reports a hardware-in-the-loop (HIL) comparison of Kolmogorov–Arnold Networks (KANs) and conventional Deep Neural Networks (DNNs) within an autonomous driving pipeline executed in closed-loop. Using a digital-twin testbed with real-time hardware execution, we evaluate driving performance, perception accuracy, and planning quality. In our experiments, KAN-based controllers achieve performance comparable to DNNs while using fewer parameters, and provide a degree of interpretability by enabling symbolic approximations of their learned functions in simplified scenarios. A hybrid KAN–DNN architecture, which integrates KAN functional layers with standard dense layers, showed both improved transparency and competitive performance across tasks. Unlike black-box DNNs, the KAN formulation permits symbolic inspection of planning policies, facilitating verification and design-time analysis. These results indicate that KANs are a viable option for embedded AI in autonomous systems, offering efficiency for resource-constrained hardware while providing opportunities for improved interpretability. Chaoran Yuan, Fadi J. Kurdahi |
DATE | 2 |
| 2025 | SoftmAP: Software-Hardware Co-Design for Integer-Only Softmax on Associative ProcessorsabstractRecent research efforts focus on reducing the computational and memory overheads of Large Language Models (LLMs) to make them feasible on resource-constrained devices. Despite advancements in compression techniques, nonlinear operators like Softmax and Layernorm remain bottlenecks due to their sensitivity to quantization. We propose SoftmAP, a software-hardware co-design methodology that implements an integer-only low-precision Softmax using In-Memory Compute (IMC) hardware. Our method achieves up to three orders of magnitude improvement in the energy-delay product compared to A100 and RTX3090 GPUs, making LLMs more deployable without compromising performance. Mariam Rakka, Jinhao Li 0006, Guohao Dai 0001, Ahmed M. Eltawil, Mohamed E. Fouda, Fadi J. Kurdahi |
DATE | 6 |
| 2025 | GDS2SEM: Diffusion-based Layout-to-SEM Post-Fabrication Emulation for IC ValidationabstractContinued technology scaling and ever shrinking feature sizes in Integrated Circuit (IC) layouts cause those layouts to look like highly idealized versions of the true structures on a manufactured IC. However, there are a number of instances where we need to be able to link the ideal and manufactured layouts like for quality control and reverse engineering. The imaging of ICs is performed via Scanning Electron Microscopy (SEM). It is a complex and expensive process which is invaluable when debugging technology problems or finding circuit defects. An alternative is SEM or Lithography Simulation, where physics-based models of the imaging process can be used to predict the manufactured layout, which is also is highly complex and resource consuming process. Data-driven simulators also suffer from the limited amount of publicly available IC design data for training. In this paper, we present GDS2SEM, a data-driven and diffusion-based alternative to understanding the relationship between ideal and manufactured layouts, which can help us generate new Layout-SEM image data. We formulate the Layout-to-SEM mapping as a machine learning image-to-image translation task. We show that it is possible to create such an accurate mapping, and demonstrate it on a number of layouts. We also showcase the potential that this method holds for data augmentation by testing its generative capabilities on unseen data. Walaa Amer, Sani R. Nassif, Fadi J. Kurdahi |
ISCAS | 3 |
| 2025 | Mitigating the Impact of ReRAM I-V Nonlinearity and IR Drop via Fast Offline Network TrainingabstractReRAM crossbar arrays (RCAs) have the potential to provide extremely high efficiency for accelerating deep neural networks (DNNs). However, one crucial challenge for RCA-based DNN accelerators is functional inaccuracy due to nonidealities present in RCA hardware. While nonideality-aware training (NAT) could be used to mitigate the effect of nonidealities, with currently available methods it would take months to train even a medium size convolutional neural network (CNN). In this article we propose a nonideality prediction method that enables very fast training of RCA-based neural networks, and show its feasibility through NAT of DNNs. Our key ideas include 1) weight-centric nonideality modeling and 2) data-dependence elimination by tailored input randomization. Our experimental results using a multilayer perceptron and CNNs demonstrate that our method is very fast ($100\sim 15$$000\times $faster training speed) while achieving much better-crossbar-level accuracy ($2 \sim 90\times $lower-RMS error) and post-retraining validated accuracy than previous methods. Sugil Lee, Mohamed E. Fouda, Chenghao Quan, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2024 | HDRLPIM: A Simulator for Hyper-Dimensional Reinforcement Learning Based on Processing In-MemoryabstractProcessing In-Memory (PIM) is a data-centric computation paradigm that performs computations inside the memory, hence eliminating the memory wall problem in traditional computational paradigms used in Von-Neumann architectures. The associative processor, a type of PIM architecture, allows performing parallel and energy-efficient operations on vectors. This architecture is found useful in vector-based applications such as Hyper-Dimensional (HDC) Reinforcement Learning (RL). HDC is rising as a new powerful and lightweight alternative to costly traditional RL models such as Deep Q-Learning. The HDC implementation of Q-Learning relies on encoding the states in a high-dimensional representation where calculating Q-values and finding the maximum one can be done entirely in parallel. In this article, we propose to implement the main operations of a HDC RL framework on the associative processor. This acceleration achieves up to \(152.3\times\) and \(6.4\times\) energy and time savings compared to an FPGA implementation. Moreover, HDRLPIM shows that an SRAM-based AP implementation promises up to \(968.2\times\) energy-delay product gains compared to the FPGA implementation. Mariam Rakka, Walaa Amer, Hanning Chen, Mohsen Imani, Fadi J. Kurdahi |
ACM J. Emerg. Technol. Comput. Syst. | 5 |
| 2024 | A Review of State-of-the-art Mixed-Precision Neural Network FrameworksabstractMixed-precision Deep Neural Networks (DNNs) provide an efficient solution for hardware deployment, especially under resource constraints, while maintaining model accuracy. Identifying the ideal bit precision for each layer, however, remains a challenge given the vast array of models, datasets, and quantization schemes, leading to an expansive search space. Recent literature has addressed this challenge, resulting in several promising frameworks. This paper offers a comprehensive overview of the standard quantization classifications prevalent in existing studies. A detailed survey of current mixed-precision frameworks is provided, with an in-depth comparative analysis highlighting their respective merits and limitations. The paper concludes with insights into potential avenues for future research in this domain. Mariam Rakka, Mohamed E. Fouda, Pramod P. Khargonekar, Fadi J. Kurdahi |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Information Processing Factory 2.0 - Self-awareness for Autonomous Collaborative SystemsabstractThis paper summarizes the talks of a special session on the IPF 2.0 project, a collaborative German-US research project that leverages self-awareness principles for the self-management of distributed systems of autonomous multiprocessor systems-on-chip (MPSoCs). Nora Sperling, Alex Bendrick, Dominik Stöhrmann, Rolf Ernst, Bryan Donyanavard, Florian Maurer 0003, Oliver Lenke, Anmol Surhonne, Andreas Herkersdorf, Walaa Amer, Caio Batista de Melo, Ping-Xiang Chen, Quang Anh Hoang, Rachid Karami, Biswadip Maity, Paul Nikolian, Mariam Rakka, Dongjoo Seo, Saehanseul Yi, Minjun Seo, Nikil Dutt, Fadi J. Kurdahi |
DATE | 22 |
| 2023 | Hardware Implementation and Evaluation of an Information Processing FactoryabstractThe Information Processing Factory (IPF) utilizes factory management principles to tackle the complexities of integrated embedded systems, ensuring continuous safe operation and optimization at runtime. This paper presents a hardware implementation of IPF that enables dynamic task migration across system resources, ensuring reliability in the face of internal or external failures. We demonstrate the effectiveness of IPF through the efficient migration of tasks in multiprocessor SoCs using a safety-critical pacemaker application as a case study. Despite the additional software and hardware requirements, implementing IPF in a pacemaker results in comparable reliability to dual modular redundancy (DMR) with faster service resumption and improved resource utilization. Walaa Amer, Mariam Rakka, Rachid Karami, Minjun Seo, Mazen A. R. Saghir, Rouwaida Kanj, Fadi J. Kurdahi |
VLSI-SoC | 7 |
| 2023 | An Accurate Non-accelerometer-based PPG Motion Artifact Removal Technique using CycleGANabstractA photoplethysmography (PPG) is an uncomplicated and inexpensive optical technique widely used in the healthcare domain to extract valuable health-related information, e.g., heart rate variability, blood pressure, and respiration rate. PPG signals can easily be collected continuously and remotely using portable wearable devices. However, these measuring devices are vulnerable to motion artifacts caused by daily life activities. The most common ways to eliminate motion artifacts use extra accelerometer sensors, which suffer from two limitations: (i) high power consumption, and (ii) the need to integrate an accelerometer sensor in a wearable device (which is not required in certain wearables). This paper proposes a low-power non-accelerometer-based PPG motion artifacts removal method outperforming the accuracy of the existing methods. We use Cycle Generative Adversarial Network to reconstruct clean PPG signals from noisy PPG signals. Our novel machine-learning-based technique achieves 9.5 times improvement in motion artifact removal compared to the state-of-the-art without using extra sensors such as an accelerometer, which leads to 45% improvement in energy efficiency. Amir Hosein Afandizadeh Zargari, Seyed Amir Hossein Aqajari, Hadi Khodabandeh, Amir-Mohammad Rahmani, Fadi J. Kurdahi |
ACM Trans. Comput. Heal. | 5 |
| 2023 | Resistive Neural Hardware AcceleratorsabstractDeep neural networks (DNNs), as a subset of machine learning (ML) techniques, entail that real-world data can be learned, and decisions can be made in real time. However, their wide adoption is hindered by a number of software and hardware limitations. The existing general-purpose hardware platforms used to accelerate DNNs are facing new challenges associated with the growing amount of data and are exponentially increasing the complexity of computations. Emerging nonvolatile memory (NVM) devices and the compute-in-memory (CIM) paradigm are creating a new hardware architecture generation with increased computing and storage capabilities. In particular, the shift toward resistive random access memory (ReRAM)-based in-memory computing has great potential in the implementation of area- and power-efficient inference and in training large-scale neural network architectures. These can accelerate the process of IoT-enabled AI technologies entering our daily lives. In this survey, we review the state-of-the-art ReRAM-based DNN many-core accelerators, and their superiority compared to CMOS counterparts was shown. The review covers different aspects of hardware and software realization of DNN accelerators, their present limitations, and prospects. In particular, a comparison of the accelerators shows the need for the introduction of new performance metrics and benchmarking standards. In addition, the major concerns regarding the efficient design of accelerators include a lack of accuracy in simulation tools for software and hardware codesign. Kamilya Smagulova, Mohamed E. Fouda, Fadi J. Kurdahi, Khaled N. Salama, Ahmed M. Eltawil |
Proc. IEEE | 3 |
| 2023 | Offline Training-Based Mitigation of IR Drop for ReRAM-Based Deep Neural Network AcceleratorsabstractRecently, resistive RAM (ReRAM)-based hardware accelerators showed unprecedented performance compared the digital accelerators. Technology scaling causes an inevitable increase in interconnect wire resistance, which leads to IR drops that could limit the performance of ReRAM-based accelerators. These IR drops deteriorate the signal integrity and quality, especially in the crossbar structures which are used to build high-density ReRAMs. Hence, finding a software solution, which can predict the effect of IR drop without involving expensive hardware or SPICE simulations, is very desirable. In this article, we propose two neural networks models to predict the impact of the IR drop problem. These models are used to evaluate the performance of the different deep neural network (DNN) models including binary and quantized neural networks showing similar performance (i.e., recognition accuracy) to the golden validation (i.e., SPICE-based DNN validation). In addition, these predication models are incorporated into the DNN training framework to efficiently retrain the DNN models and bridge the accuracy drop. To further enhance the validation accuracy, we propose incremental training methods. The DNN validation results, done through SPICE simulations, show very high improvement in performance close to the baseline performance, which demonstrates the efficacy of the proposed method even with challenging datasets, such as CIFAR10 and SVHN. Sugil Lee, Mohamed E. Fouda, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2023 | Training-Free Stuck-At Fault Mitigation for ReRAM-Based Deep Learning AcceleratorsabstractAlthough Resistive RAMs can support highly efficient matrix–vector multiplication, which is very useful for machine learning and other applications, the nonideal behavior of hardware, such as stuck-at fault (SAF) and IR drop is an important concern in making ReRAM crossbar array-based deep learning accelerators. Previous work has addressed the nonideality problem through either redundancy in hardware, which requires a permanent increase of hardware cost, or software retraining, which may be even more costly or unacceptable due to its need for a training dataset as well as high computation overhead. In this article, we propose a very lightweight method that can be applied on top of existing hardware or software solutions. Our method, called forward-parameter tuning (FPT), takes advantage of a certain statistical property existing in the activation data of neural network layers, and can mitigate the impact of mild nonidealities in ReRAM crossbar arrays (RCAs) for deep learning applications without using any hardware, a dataset, or gradient-based training. Our experimental results using MNIST, CIFAR-10, and CIFAR-100, and ImageNet datasets in binary and multibit networks demonstrate that our technique is very effective, both alone and together with previous methods, up to 20% fault rate, which is higher than even some of the previous remapping methods. We also evaluate our method in the presence of other nonidealities, such as variability and IR drop. Furthermore, we provide an analysis based on the concept of the effective fault rate (EFR), which not only demonstrates that EFR can be a useful tool to predict the accuracy of faulty RCA-based neural networks but also explains why mitigating the SAF problem is more difficult with multibit neural networks. Chenghao Quan, Mohamed E. Fouda, Sugil Lee, Giju Jung, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2022 | Accurate Prediction of ReRAM Crossbar Performance Under I-V Nonlinearity and IR DropabstractDespite the promise of extremely efficient matrix-vector multiplication (MVM) by ReRAM crossbar arrays (RCAs), maintaining high accuracy has been challenging due to nonidealities such as wire resistance (also known as IR drop) and I-V nonlinearity (i.e., voltage-dependent conductance). For system architects, a fast method to accurately predict the MVM output of an RCA under nonidealities is highly desirable. While IR drop alone without I-V nonlinearity can be efficiently predicted, the existence of I-V nonlinearity makes the problem much harder. In this paper we propose a novel algorithm based on iterative refinement, which can predict with high accuracy the outcome of an MVM operation on an RCA in the presence of both I-V nonlinearity and IR drop. Our experiments using binary RCAs of different sizes demonstrate that our proposed method is order-of-magnitude more accurate than previous methods in terms of RMS error. We also present case studies predicting hardware-realistic accuracy of binarized neural networks on RCAs as well as nonideality-aware retraining, demonstrating the efficacy of our method for early design space exploration of ReRAM-based accelerators. Sugil Lee, Mohamed E. Fouda, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi |
ICCD | 5 |
| 2022 | Guest Editorial: Secure Radio-Frequency (RF)-Analog Electronics and ElectromagneticsabstractNo abstract available. Vanessa Chen, Mohammad Al Faruque, Fadi J. Kurdahi |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2022 | CFFNN: Cross Feature Fusion Neural Network for Collaborative FilteringabstractNumerous state-of-the-art recommendation frameworks employ deep neural networks in Collaborative Filtering (CF). In this paper, we propose a cross feature fusion neural network (CFFNN) for the enhancement of CF. Existing studies overlook either user preferences for various item features or the relationship between item features and user features. To solve this problem, we construct a cross feature fusion network to enable the fusion of user features and item features as well as a self-attention network to determine users’ preferences for items. Specifically, we design a feature extraction layer with multiple MLP (Multilayer Perceptrons) modules to extract both user features and item features. Then, we introduce a cross feature fusion mechanism for an accurate determination of the relationship between different user-item interactions. The features of users and items are crossly embedded and then fed into a prediction network. The attention mechanism enables the model to focus on more effective features. The effectiveness of CFFNN model is demonstrated through extensive experiments on four real-world datasets. The experimental results indicate that CFFNN significantly outperforms the existing state-of-the-art models, with a relative improvement of 3.0 to 12.1 percent on hit ratio (HR) and normalized discounted cumulative gain (NDCG) compared with the baselines. Ruiyun Yu, Dezhi Ye, Biyun Zhang, Ann Move Oguti, Jie Li 0008, Bo Jin 0001, Fadi J. Kurdahi |
IEEE Trans. Knowl. Data Eng. | 8 |
| 2021 | Cost- and Dataset-free Stuck-at Fault Mitigation for ReRAM-based Deep Learning AcceleratorsabstractResistive RAMs can implement extremely efficient matrix vector multiplication, drawing much attention for deep learning accelerator research. However, high fault rate is one of the fundamental challenges of ReRAM crossbar array-based deep learning accelerators. In this paper we propose a dataset-free, cost-free method to mitigate the impact of stuck-at faults in ReRAM crossbar arrays for deep learning applications. Our technique exploits the statistical properties of deep learning applications, hence complementary to previous hardware or algorithmic methods. Our experimental results using MNIST and CIFAR-10 datasets in binary networks demonstrate that our technique is very effective, both alone and together with previous methods, up to 20 % fault rate, which is higher than the previous remapping methods. We also evaluate our method in the presence of other non-idealities such as variability and IR drop. Giju Jung, Mohamed E. Fouda, Sugil Lee, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi |
DATE | 6 |
| 2021 | Fast and Low-Cost Mitigation of ReRAM Variability for Deep Learning ApplicationsabstractTo overcome the programming variability (PV) of ReRAM crossbar arrays (RCAs), the most common method is program-verify, which, however, has high energy and latency overhead. In this paper we propose a very fast and low-cost method to mitigate the effect of PV and other variability for RCA-based DNN (Deep Neural Network) accelerators. Leveraging the statistical properties of DNN output, our method called Online Batch-Norm Correction (OBNC) can compensate for the effect of programming and other variability on RCA output without using on-chip training or an iterative procedure, and is thus very fast. Also our method does not require a nonideality model or a training dataset, hence very easy to apply. Our experimental results using ternary neural networks with binary and 4-bit activations demonstrate that our OBNC can recover the baseline performance in many variability settings and that our method outperforms a previously known method (VCAM) by large margins when input distribution is asymmetric or activation is multi-bit. Sugil Lee, Mohamed E. Fouda, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi |
ICCD | 5 |
| 2020 | Learning to Predict IR Drop with Effective Training for ReRAM-based Neural Network HardwareabstractDue to the inevitability of the IR drop problem in passive ReRAM crossbar arrays, finding a software solution that can predict the effect of IR drop without the need of expensive SPICE simulations, is very desirable. In this paper, two simple neural networks are proposed as software solution to predict the effect of IR drop. These networks can be easily integrated in any deep neural network framework to incorporate the IR drop problem during training. As an example, the proposed solution is integrated in BinaryNet framework and the test validation results, done through SPICE simulations, show very high improvement in performance close to the baseline performance, which demonstrates the efficacy of the proposed method. In addition, the proposed solution outperforms the prior work on challenging datasets such as CIFAR10 and SVHN. Sugil Lee, Giju Jung, Mohamed E. Fouda, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi |
DAC | 6 |
| 2020 | Towards Self-Aware Systems-on-Chip Through Intelligent Cross-Layer CoordinationabstractAlthough there is a rich history of cross-layer design for embedded computing systems to achieve desired QoS, we are facing ever more challenges from the intertwined goals of energy- efficiency, thermal design constraints, as well as resilience to errors emanating from the application, environment and hardware platforms. We posit that next-generation computing platforms must necessarily deploy intelligent cross-layer design achieved through self-awareness principles inspired by biology and nature. Such an approach will move us from current strategies (using limited cross-layer coordination) to a holistic cross-layer strategy that enables intelligent cross-layer management policies which can adaptively tune itself based on the current state of the system. The talk will present design exemplars that embrace this intelligent cross-layer approach, and highlight the role of self-awareness in achieving dynamic adaptivity. Fadi J. Kurdahi |
ACM Great Lakes Symposium on VLSI | 1 |
| 2020 | NEWERTRACK: ML-Based Accurate Tracking of In-Mouth Nutrient Sensors Position Using Spectrum-Wide InformationabstractIn this article, we consider a family of in-mouth passive radio-frequency (RF) sensors for nutrients whose data is acquired externally using a miniaturized vector network analyzer (VNA). However, the data readings suffer from noise as the VNA's position shifts during reading and also in-between readings. We propose an ML method to track the movement of these sensors using RF signals. Although classical studies in RF sensors have shown that the s-parameter of the sensor is related to the position of the sensor, they have always tried to find the relation between the resonance frequency and its loss to the position of the sensor. In contrast, our analysis revealed that based on theoretical perspectives, using just resonance frequency and its loss makes it impossible to find the correct position of the sensor with reasonable accuracy. To improve accuracy, we propose to use the loss information of RF sensors not only at one resonance frequency but across a wide spectrum. We introduced a pretrained neural network model consisting of a combination of DNN and convolutional neural network layers for extracting the sensor position. In order to accurately track the movement of the sensor with respect to the VNA, we propose a neural network-based RF sensor tracking (NEWERTRACK) model, which has utilized two unidirectional long short term memory layers. The results show that our proposed model achieves an average tracking error of 1.84 mm. Also, NEWERTRACK is capable of improving the position tracking accuracy by about 10× compared to traditional s-parameter analysis methods and outperforms the best model in other RF methods by two times. Amir Hosein Afandizadeh Zargari, Manik Dautta, Marzieh Ashrafiamiri, Minjun Seo, Peter Tseng, Fadi J. Kurdahi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2019 | Non-Stationary Polar Codes for Resistive MemoriesabstractResistive memories are considered a promising memory technology enabling high storage densities. However, the readout reliability of resistive memories is impaired due to the inevitable existence of wire resistance, resulting in the sneak path problem. Motivated by this problem, we study polar coding over channels with different reliability levels, termed non-stationary polar codes, and we propose a technique improving the bit error rate (BER) performance. We then apply the framework of non-stationary polar codes to the crossbar array and evaluate its BER performance under two modeling approaches, namely binary symmetric channels and binary asymmetric channels. Finally, we propose a technique for biasing the proportion of high-resistance states in the crossbar array and show its advantage in reducing further the BER. Several simulations are carried out using a SPICE-like simulator, exhibiting significant reduction in BER. Marwen Zorgui, Mohamed E. Fouda, Zhiying Wang 0001, Ahmed M. Eltawil, Fadi J. Kurdahi |
GLOBECOM | 5 |
| 2019 | Testing Topology Adaptive Irrigation IoT with CircuitsabstractThere is a significant unrealized potential in developing state of the art electronics for agriculture, particularly, for irrigation systems. However, development velocity in the domain of Irrigation Internet of Things (IrIoT) is slowed due to lengthy and complex validation cycles that require multi domain integration testing. This paper proposes testing IrIoT distributed controllers on electrical circuit platform prior to final verification. This flexible testing environment utilizes wires as opposed to water lines for testing design shortcomings. To fully expose challenges associated with next generation IrIoT this paper discusses testing challenges of topology adaptive distributed wireless irrigation controllers. Here are presented contributions in topology adaptation method, software simulation tools, an intermediate testing step using circuits as opposed to directly conducting integration testing in the target setting. The proposed method establishes an isolated circuit sandbox where integral system components, in this case IrIoT controllers, and algorithms are tested prior to full integration testing. Davit Hovhannisyan, Ahmed M. Eltawil, Fadi J. Kurdahi |
ISCAS | 3 |
| 2019 | Feasibility Study of Plant Health MonitoringabstractContinuous monitoring of crop is an essential task of agricultural practices for detection of diseases or pests, precision irrigation and fertilization. The state of the art monitoring and imaging systems use aerial imaging to obtain visual feedback and multi-spectral imagery to determine crop growth factors. The main idea is that the features can be automatically calculated and assessed after pre-processing the images. After pre-processing, then the parameters can be computed using image processing techniques. For example, key leaf function traits leaf life span, leaf mass per area can be calculated. Our findings indicate that plant health assessment could be moved from lab and expensive monitoring tools to ubiquitous silicon technology based cost effect solutions without much loss of accuracy. Davit Hovhannisyan, Kareem Khalifeh, Peng Fei, Ahmed M. Eltawil, Fadi J. Kurdahi |
ISCAS | 5 |
| 2019 | Hybrid pyramid-DWT-SVD dual data hiding technique for videos ownership protection
Farhan A. Alenizi, Fadi J. Kurdahi, Ahmed M. Eltawil, Awad Kh. Al-Asmari |
Multim. Tools Appl. | 2 |
| 2019 | Efficient Tracing Methodology Using Automata ProcessorabstractTracing or trace interface has been used in various ways to find system defects or bugs. As embedded systems are increasingly used in safety-critical applications, tracing can provide useful information during system execution at runtime. Non-intrusive tracing that does not affect system performance has become especially important, but unfortunately, the biggest obstacle to this approach was the vast amount of real-time trace data, making it challenging to address complex requirements with relatively limited hardware implementations. Automata processors can be programmed with a memory-like structure of automata and have a structure specific to streaming data, large capacity, and parallel processing functions. This paper promotes the idea of high-level system-on-chip monitoring using automata processors. We used a safety-critical pacemaker application in the experiments, described timed automata (TA)-based requirements, and tested intentionally injected 4,000 random failures. The TA model converted for Automata Processor to monitor system, correctness, and safety properties achieved 100% failure detection rate in the experiment, and the detected failure is reported as fast enough to allow enough extent for failure recovery. Minjun Seo, Fadi J. Kurdahi |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2018 | Rapid in-memory matrix multiplication using associative processorabstractMemory hierarchy latency is one of the main problems that prevents processors from achieving high performance. To eliminate the need of loading/storing large sets of data, Resistive Associative Processors (ReAP) have been proposed as a solution to the von Neumann bottleneck. In ReAPs, logic and memory structures are combined together to allow inmemory computations. In this paper, we propose a new algorithm to compute the matrix multiplication inside the memory that exploits the benefits of ReAP. The proposed approach is based on the Cannon algorithm and uses a series of rotations without duplicating the data. It runs in O(n), where n is the dimension of the matrix. The method also applies to a large set of row by column matrix-based applications. Experimental results show several orders of magnitude increase in performance and reduction in energy and area when compared to the latest FPGA and CPU implementations. Mohamed A. Neggaz, Hasan Erdem Yantir, Smaïl Niar, Ahmed M. Eltawil, Fadi J. Kurdahi |
DATE | 5 |
| 2018 | Design methodologies for enabling self-awareness in autonomous systemsabstractThis paper deals with challenges and possible solutions for incorporating self-awareness principles in EDA design flows for autonomous systems. We present a holistic approach that enables self-awareness across the software/hardware stack, from systems-on-chip to systems-of-systems (autonomous car) contexts. We use the Information Processing Factory (IPF) metaphor as an exemplar to show how self-awareness can be achieved across multiple abstraction levels, and discuss new research challenges. The IPF approach represents a paradigm shift in platform design by envisioning the move towards a consequent platform-centric design in which the combination of self-organizing learning and formal reactive methods guarantee the applicability of such cyber-physical systems in safety-critical and high-availability applications. Armin Sadighi, Bryan Donyanavard, Thawra Kadeed, Kasra Moazzemi, Tiago Rogério Mück, Ahmed Nassar 0001, Amir-Mohammad Rahmani, Thomas Wild, Nikil Dutt, Rolf Ernst, Andreas Herkersdorf, Fadi J. Kurdahi |
DATE | 12 |
| 2018 | Circuit Inspired Modeling Method for IrrigationabstractPrecision irrigation systems promise to bring significant improvement in resource efficiency and crop yield by providing analytics and smart tools for the growers. While significant amounts of data can be collected in a sensor-rich system, there are no rigorously designed models that can provide actionable intelligence to the user. This paper proposes the integration of circuit-inspired modeling of natural phenomena and man-made artifacts to generate end-to-end irrigation system circuit models. Such models can take advantage of existing circuit design and simulation tools that have been perfected over the past decades to efficiently process large input sets. We show that circuit-inspired models are indeed qualitatively sound and quantitatively accurate in capturing both natural phenomena and engineered physical irrigation systems. Davit Hovhannisyan, Ahmed M. Eltawil, Mohammad Abdullah Al Faruque, Fadi J. Kurdahi |
DSD | 4 |
| 2018 | Power optimization techniques for associative processors
Hasan Erdem Yantir, Ahmed M. Eltawil, Smaïl Niar, Fadi J. Kurdahi |
J. Syst. Archit. | 4 |
| 2018 | Platform-Centric Self-Awareness as a Key Enabler for Controlling Changes in CPSabstractFuture cyber-physical systems will host a large number of coexisting distributed applications on hardware platforms with thousands to millions of networked components communicating over open networks. These applications and networks are subject to continuous change. The current separation of design process and operation in the field will be superseded by a life-long design process of adaptation, infield integration, and update. Continuous change and evolution, application interference, environment dynamics and uncertainty lead to complex effects which must be controlled to serve a growing set of platform and application needs. Self-adaptation based on self-awareness and self-configuration has been proposed as a basis for such a continuous in-field process. Research is needed to develop automated in-field design methods and tools with the required safety, availability, and security guarantees. The paper shows two complementary use cases of self-awareness in architectures, methods, and tools for cyber-physical systems. The first use case focuses on safety and availability guarantees in self-aware vehicle platforms. It combines contracting mechanisms, tool based self-analysis and self-configuration. A software architecture and a runtime environment executing these tools and mechanisms autonomously are presented including aspects of self-protection against failures and security threats. The second use case addresses variability and long term evolution in networked MPSoC integrating hardware and software mechanisms of surveillance, monitoring, and continuous adaptation. The approach resembles the logistics and operation principles of manufacturing plants which gave rise to the metaphoric term of an Information Processing Factory that relies on incremental changes and feedback control. Both use cases are investigated by larger research groups. Despite their different approaches, both use cases face similar design and design automation challenges which will be summarized in the end. We will argue that seemingly unrelated research challenges, such as in machine learning and security, could also profit from the methods and superior modeling capabilities of self-aware systems. Mischa Möstl, Johannes Schlatow, Rolf Ernst, Nikil Dutt, Ahmed Nassar 0001, Amir-Mohammad Rahmani, Fadi J. Kurdahi, Thomas Wild, Armin Sadighi, Andreas Herkersdorf |
Proc. IEEE | 7 |
| 2018 | A Two-Dimensional Associative Processor
Hasan Erdem Yantir, Ahmed M. Eltawil, Fadi J. Kurdahi |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2017 | Efficient pulsed-latch implementation for multiport register files: work-in-progressabstractIn this paper, register file design using pulsed latches is presented. Having some advantages in performance, area and power, pulsed latches represent an attractive implementation of register files. In addition, a proposed multiport register file architecture is introduced using single physical read/write ports to virtualize additional ports for read and write. The initial results show huge savings in area and power in comparison to the traditional architectures. Wael M. Elsharkasy, Hasan Erdem Yantir, Amin Khajeh, Ahmed M. Eltawil, Fadi J. Kurdahi |
CASES | 5 |
| 2017 | Low Latency Approximate Adder for Highly Correlated Input StreamsabstractApproximate computing helps achieve better performance or energy efficiency by trading accuracy. Most approximate adders are composed of multiple sub-adders and long carry chains are split to reduce latency, thus benefiting from the fact that carry propagation across long carry chains is rare for uniformly distributed inputs. One key tradeoff of these approximate adders is between latency and error rate. The more prediction bits are used, the lower is the error rate, but the latency is longer. In this paper, we present a Correlation Aware Predictor (CAP) which utilizes spatial-temporal correlation information of input streams to predict carry-in value for sub-adders. CAP uses less prediction bits which help reduce adder latency significantly. For highly correlated input streams, we found that CAP can reduce adder latency by about 23% at the same error rate compared to prior work. We implemented a CAP-based approximate adder in Verilog and synthesized with TSMC 16nm library. Synthesis results show that CAP-based adder can reduce latency by 25% and save 13% in silicon area compared to state-of-the-art. Ahmed M. Eltawil, Fadi J. Kurdahi |
ICCD | 3 |
| 2017 | Approximate Memristive In-memory ComputingabstractThe bottleneck between the processing elements and memory is the biggest issue contributing to the scalability problem in computing. In-memory computation is an alternative approach that combines memory and processor in the same location, and eliminates the potential memory bottlenecks. Associative processors are a promising candidate for in-memory computation, however the existing implementations have been deemed too costly and power hungry. Approximate computing is another promising approach for energy-efficient digital system designs where it sacrifices the accuracy for the sake of energy reduction and speedup in error-resilient applications. In this study, approximate in-memory computing is introduced in memristive associative processors. Two approximate computing methodologies are proposed; bit trimming and memristance scaling. Results show that the proposed methods not only reduce energy consumption of in-memory parallel computing but also improve their performance. As compared to other existing approximate computing methodologies on different architectures (e.g., CPU, GPU, and ASIC), approximate memristive in-memory computing exhibits better results in terms of energy reduction (up to 80x) and speedup (up to 20x) on a variety of benchmarks from different domains when quality degradation is limited to 10% and it confirms that memristive associative processors provide a highly-promising platform for approximate computing. Hasan Erdem Yantir, Ahmed M. Eltawil, Fadi J. Kurdahi |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2016 | Lattice-based Boolean diagramsabstractThis paper presents lattice-based Boolean diagrams (LBBDs), a graphical representation of Boolean functions that is not derived from binary decision diagrams (BDDs), as well as symbolic manipulation algorithms. It also identifies a class of Boolean functions where LBBDs are demonstrably more efficient to construct, and reason with, when compared to BDDs. The case studies include ITC99 and MCNC benchmarks, randomly generated cube covers or sum-of-products (SOP) formulas as well as multi-level Boolean formulas. Finally, LBBDs proved to be instrumental to the efficient runtime verification of software over distributed multiprocessor systems. Ahmed Nassar 0001, Fadi J. Kurdahi |
ASP-DAC | 2 |
| 2016 | Topaz: Mining high-level safety properties from logic simulation traces
Ahmed Nassar 0001, Fadi J. Kurdahi, Salam R. Zantout |
DATE | 2 |
| 2016 | Process variations-aware resistive associative processor designabstractRecent breakthroughs in memristive devices have demonstrated the potential of using resistive content addressable memories for associative processing. These architectures enable ultra-high density integrated circuits along with low-power computation. However, the reliability of memristive elements is limiting the widespread adoption of these architectures. In this study, we address the reliability issues that arise in high density, resistive associative processor architectures. We propose methods to design process variation immune resistive content addressable memories and minimize the error probabilities. According to SPICE-based circuit simulations, the reliability of the circuit increases significantly and thus positively influences the accuracy of arithmetic operations as well. Hasan Erdem Yantir, Mohamed E. Fouda, Ahmed M. Eltawil, Fadi J. Kurdahi |
ICCD | 4 |
| 2015 | NUVA: Architectural support for runtime verification of parametric specifications over multicoresabstractRuntime Verification (RV) has recently emerged as a complementary technology to extend coverage of conventional software verification methods. To address the substantial performance and power overhead of pure software RV frameworks, this paper introduces NUVA, which stands for nonuniform verification architecture, a distributed automata-based RV architecture for parametric specifications in the form of parameterized finite-state automata, with a case study over a cache-coherent nonuniform-memory-access (ccNUMA) multiprocessor. The core of NUVA is a coherent distributed automata transactional memory (ATM) that efficiently maintains states of a dynamic population of automata checkers organized into a rooted dynamic directed acyclic graph (DAG) concurrently shared among all processor nodes. A cycle-accurate model of a ccNUMA multiprocessor confirms that performance slowdown is 1~3% and NoC message traffic increases by 10~15% at parametric event density1of 0.025 EPI for two compute-intensive scientific benchmarks having irregular concurrent data structures. The detailed architecture and implementation metrics in TSMC 40nm CMOS technology are presented. NUVA can be dimensioned to incur average total power overhead of less than 140mW and area overhead of 4mm2for a quad-core multiprocessor chip, at operating frequency of 250MHz. It is also estimated to incur 1.9~2.6% area overhead and 0.5~1% power overhead when integrated with Intel's family of high-end desktop and mobile quad-core processors operating at 1.6~3.2GHz. Our silicon implementation achieves average performance2of 1.5 MEPS/mW. For a processor core with 5 MIPS/mW, this corresponds to a 3.2% drop in energy efficiency at parametric event density of 0.01 EPI. Ahmed Nassar 0001, Fadi J. Kurdahi, Wael M. Elsharkasy |
CASES | 2 |
| 2014 | State dependent statistical timing model for voltage scaled circuitsabstractThis paper presents a novel statistical state-dependent timing model for voltage over scaled (VoS) logic circuits that accurately and rapidly finds the timing distribution of output bits. Using this model erroneous VoS circuits can be represented as error-free circuits combined with an error-injector. A case study of a two point DFT unit employing the proposed model is presented and compared to HSPICE circuit simulation. Results show an accurate match, with significant speedup gains. Aras Pirbadian, Muhammed S. Khairy, Ahmed M. Eltawil, Fadi J. Kurdahi |
ISCAS | 4 |
| 2014 | Low power reduced-complexity error-resilient MIMO detectorabstractThis paper presents a reduced-complexity low power error-resilient K-Best MIMO Detector. A novel tree-enumeration method is proposed such that the error-resilient detection processes a reduced search space and is more suitable for VLSI design. Moreover, a circuit-level optimization is employed to further simplify the complexity. Experimental results are given showing that the circuit-level optimization decreases the detector area by 15% and power consumption by 41%. Moreover, we show that the proposed error-resilient MIMO detector with reduced-voltage memory can achieve a total of 19% reduction in power consumption compared with the conventional scheme, while still maintaining close-to optimal PER performance. Chung-An Shen, Muhammed S. Khairy, Ahmed M. Eltawil, Fadi J. Kurdahi |
ISCAS | 4 |
| 2014 | Mobile Collaborative VideoabstractThe emergence of pico projectors as a part of future mobile devices presents unique opportunities for collaborative settings, especially in entertainment applications, such as video playback. By aggregating pico projectors from several users, it is possible to enhance resolution, brightness, or frame rate. In this paper we present a camera-based methodology for the alignment and synchronization of multiple projectors. The approach does not require any complicated ad hoc network setup among the mobile devices. A prototype system has been set up and used to test the proposed techniques. Kiarash Amiri, Shih-Hsien Yang, Aditi Majumder, Fadi J. Kurdahi, Magda El Zarki |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2014 | Multicopy Cache: A Highly Energy-Efficient Cache ArchitectureabstractCaches are known to consume a large part of total microprocessor energy. Traditionally, voltage scaling has been used to reduce both dynamic and leakage power in caches. However, aggressive voltage reduction causes process-variation-induced failures in cache SRAM arrays, thus compromising cache reliability. We present MultiCopy Cache (MC 2 ), a new cache architecture that achieves significant reduction in energy consumption through aggressive voltage scaling while maintaining high error resilience (reliability) by exploiting multiple copies of each data item in the cache. Unlike many previous approaches, MC 2 does not require any error map characterization and therefore is responsive to changing operating conditions (e.g., Vdd noise, temperature, and leakage) of the cache. MC 2 also incurs significantly lower overheads compared to other ECC-based caches. Our experimental results on embedded benchmarks demonstrate that MC 2 achieves up to 60% reduction in energy and energy-delay product (EDP) with only 3.5% reduction in IPC and no appreciable area overhead. Arup Chakraborty, Houman Homayoun, Amin Khajeh, Nikil Dutt, Ahmed M. Eltawil, Fadi J. Kurdahi |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2013 | Heterogeneous memory management for 3D-DRAM and external DRAM with QoSabstractThis paper presents an innovative memory management approach to utilize both 3D-DRAM and external DRAM (ex-DRAM). Our approach dynamically allocates and relocates memory blocks between the 3D-DRAM and the ex-DRAM to exploit the high memory bandwidth and the low memory latency of the 3D-DRAM as well as the high capacity and the low cost of the ex-DRAM. Our simulation shows that in workloads that are not memory intensive, our memory management technique transfers all active memory blocks to the 3D-DRAM which runs faster than the ex-DRAM. In memory intensive workloads, our memory management technique utilizes both the 3D-DRAM and the ex-DRAM to increase the memory bandwidth to alleviate bandwidth congestion. Our approach supports Quality of Service (QoS) for “latency sensitive”, “bandwidth sensitive”, and “insensitive” applications. To improve the performance and satisfy a certain level of QoS, memory blocks of different application types are allocated differently. Compared to the scratchpad memory management mechanism, the average memory access latency of our approach decreases by 19% and 23%, while performance improves by up to 5% and 12% in single threaded benchmarks and multi-threaded benchmarks respectively. Moreover, using our approach, applications do not need to manage memory explicitly like in the scratchpad case. Our memory block relocation comes with negligible performance overhead, particularly for applications which have high spatial memory locality. Le-Nguyen Tran, Fadi J. Kurdahi, Ahmed M. Eltawil, Houman Homayoun |
ASP-DAC | 2 |
| 2012 | Error resilient MIMO detector for memory-dominated wireless communication systemsabstractIn current broadband MIMO-OFDM systems such as 3GPP LTE, embedded buffering memories occupy a large portion of chip area and a significant amount of power consumption. Due to the dense structure of memories, they are especially vulnerable to scaling effects such as process variation. These effects (hardware errors) become more pronounced when aggressive voltage scaling is used due to the reduced voltage overhead. To address this issue, we present an error resilient MIMO detector. First, we derive a combined distribution of the received data in a MIMO-OFDM receiver that includes both the noise incurred by the wireless channel and errors introduced at the receiver buffering memory due to aggressive voltage scaling. Using the derived distribution, a modified MIMO detection algorithm based on the tree-searching structure is presented. A case study is presented showing that the proposed approach can achieve near-optimal performance in the presence of both channel noise and memory error, while 40% to 50% of memory power savings are realized. Muhammed S. Khairy, Chung-An Shen, Ahmed M. Eltawil, Fadi J. Kurdahi |
GLOBECOM | 4 |
| 2012 | Fast error aware model for arithmetic and logic circuitsabstractAs a result of supply voltage reduction and process variations effects, the error free margin for dynamic voltage scaling has been drastically reduced. This paper presents an error aware model for arithmetic and logic circuits that accurately and rapidly estimates the propagation delays of the output bits in a digital block operating under voltage scaling to identify circuit-level failures (timing violations) within the block. Consequently, these failure models are then used to examine how circuit-level failures affect system-level reliability. A case study consisting of a CORDIC DSP unit employing the proposed model provides tradeoffs between power, performance and reliability. Samy Zaynoun, Muhammed S. Khairy, Ahmed M. Eltawil, Fadi J. Kurdahi, Amin Khajeh |
ICCD | 4 |
| 2012 | Collaborative video playback on a federation of tiled mobile projectors enabled by visual feedbackabstractPico projectors are expected to become increasingly popular in the near future, in particular when embedded in mobile devices such as smart phones, portable media players and digital cameras. Consumers seeking more features and attracted to all in one portable devices, will find such integrated devices very appealing. However the resolution and the brightness provided by integrated mobile projectors are much lower than what standard projectors commonly offer. Our collaborative scheme based on our synchronization technique provides a means of increasing resolution and brightness for the video projected by such mobile devices, significantly enhancing the viewing experience for the user. Kiarash Amiri, Shih-Hsien Yang, Fadi J. Kurdahi, Magda El Zarki, Aditi Majumder |
MMSys | 3 |
| 2012 | Parity-based mono-Copy Cache for low power consumption and high reliabilityabstractThe power consumption is one of the most important preoccupations of the chip designers. However, reducing power consumption has its negative impact on the circuit. For example, reducing the supply voltage of a microprocessor implies an increase in the probability of process-variation-induced failures. Fault tolerant architectures propose a trade-off by boosting the reliability while reducing power consumption. Since a large part of the microprocessor power is consumed by the cache memory, we propose in this paper the Parity-based mono-Copy Cache (PmC2) that maintains cache reliability under aggressive voltage scaling. PmC2results in reducing energy consumption considerably with very low performance penalty. PmC2uses a parity check mechanism in error detection and only one cache block redundancy for error correction. Our experimental results demonstrate that reducing the supply voltage with roughly 25% of nominal Vdd achieves more than 62% reduction in cache power consumption with a negligible IPC loss that does not exceed 0.15%. Ihsen Alouani, Smaïl Niar, Fadi J. Kurdahi, Mohamed Abid |
RSP | 3 |
| 2012 | Error-Aware Algorithm/Architecture Coexploration for Video Over Wireless ApplicationsabstractIn this article, we propose a cross-layer algorithm/architecture coexploration for wireless multimedia systems to coordinate interactions among sublayer optimizers for improvements in energy/QoS/reliability. By exploiting the inherent redundancy in wireless multimedia systems, we generate an expanded design space over traditional layer-specific approaches. Specifically, we control the error resilient encoder at the application layer to provide awareness of architectural exploration at the physical layer allowing new design points with lower power consumption via aggressive voltage scaling. While trying to reduce energy consumption, the fault tolerant technique compensates the effect of the hardware and network errors due to aggressive voltage scaling and lossy transmission, respectively. Our experiments on H.263 video over a WCDMA communication system demonstrate that coexploration enlarges the feasible design space, which results in significant power savings of more than 20% in the WCDMA modem. Amin Khajeh, Minyoung Kim 0002, Nikil Dutt, Ahmed M. Eltawil, Fadi J. Kurdahi |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2012 | Variation Trained Drowsy Cache (VTD-Cache): A History Trained Variation Aware Drowsy Cache for Fine Grain Voltage ScalingabstractIn this paper we present the “Variation Trained Drowsy Cache” (VTD-Cache) architecture. VTD-Cache allows for a significant reduction in power consumption while addressing reliability issues raised by memory cell process variability. By managing voltage scaling at a very fine granularity, each cache way can be sourced at a different voltage where the selection of voltage levels depends on both the vulnerability of the memory cells in that cache way to process variation and the likelihood of access to that cache location. After a short training period, the proposed architecture will micro-tune the cache, allowing significant power reduction with negligible increase in the number of misses. In addition, the proposed architecture actively monitors the access pattern and reconfigures the supply voltage setting to adapt to the execution pattern of the program. The novel and modular architecture of the VTD-Cache and its associated controller makes it easy to be implemented in memory compilers with a small area and power overhead. In a case study, the SimpleScalar simulation of the proposed 32 kB cache architecture reports over 57% reduction in power consumption over standard SPEC2000 integer benchmarks while incurring an area overhead of less than 4% and an execution time penalty smaller than 1%. Avesta Sasan, Kiarash Amiri, Houman Homayoun, Ahmed M. Eltawil, Fadi J. Kurdahi |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2011 | A Class of Low Power Error Compensation Iterative DecodersabstractRecent power reduction techniques aggressively modulate the supply voltage of embedded buffering memories allowing acceptable hardware errors to flow through the processing chain. In this paper, we introduce a class of modified Turbo and LDPC decoders that provide significant improvements over standard decoders in the presence of hardware noise. Simulation results show a consistent improvement in the BER performance of the modified decoders across all SNRs with very small area and power overheads as compared to the conventional decoders. Amr M. A. Hussien, Muhammed S. Khairy, Amin Khajeh, Ahmed M. Eltawil, Fadi J. Kurdahi |
GLOBECOM | 5 |
| 2011 | Reconfigurable filter implementation of a matched-filter based spectrum sensor for Cognitive Radio systemsabstractSpectrum sensing is one of the most important features of Cognitive Radio (CR) systems. Matched-filter based spectrum sensing techniques provide optimum sensing performance given that a number of characteristics of the transmitted signal are known by the sensors. Assuming that the received signal pertains to one communication standard from a given set of wireless technologies, conventional spectrum sensors employ separate filters corresponding to each standard which gives rise to increased power consumption and ciruit size. A novel reconfigurable matched-filter based spectrum sensor to be deployed in CR systems is proposed in order to overcome the disadvantages of conventional design methods. This approach proposes a spectrum of design qualities which trade-off area for reconfiguration overhead. We will show that our approach is capable of designing reconfigurable filter for standards with widely varying filter characteristics. Amir Hossein Gholamipour, Ali Gorcin, B. Ugur Töreyin, Mazen A. R. Saghir, Fadi J. Kurdahi, Ahmed M. Eltawil |
ISCAS | 6 |
| 2011 | Embedded Memories Fault-Tolerant Pre- and Post-Silicon OptimizationabstractThis paper proposes a structured method for scaling both the supply voltage as well as the body bias voltage for CMOS embedded static memory with the aim of achieving a controllable and dynamic probability of failure with minimum power consumption for each memory block. The target error probability is managed according to the time varying error tolerance attributes of the application using the memory at a certain instant in time. This approach enables system designers to abstract the concepts of power awareness, yield and reliability as design tradeoffs-that incorporate application knowledge-early in the design cycle. The paper develops a formal theoretical and practical foundation based on the underlying device statistics upon which both system and circuit designers can investigate error aware design. Amin Khajeh, Ahmed M. Eltawil, Fadi J. Kurdahi |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2011 | A Multi-Granularity Power Modeling Methodology for Embedded ProcessorsabstractWith power becoming a major constraint for multiprocessor embedded systems, it is becoming important for designers to characterize and model processor power dissipation. It is critical for these processor power models to be useable across various modeling abstractions in an electronic system level (ESL) design flow, to guide early design decisions. In this paper, we propose a unified processor power modeling methodology for the creation of power models at multiple granularity levels that can be quickly mapped to an ESL design flow. Our experimental results based on applying the proposed methodology on the OpenRISC and MIPS processors demonstrate the usefulness of having multiple power models. The generated models range from very high-level two-state and architectural/instruction set simulator models that can be used in transaction level models, to extremely detailed cycle-accurate models that enable early exploration of power optimization techniques. These models offer a designer tremendous flexibility to trade off estimation accuracy with estimation/simulation effort. Young-Hwan Park, Sudeep Pasricha, Fadi J. Kurdahi, Nikil Dutt |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2011 | Inquisitive Defect Cache: A Means of Combating Manufacturing Induced Process VariationabstractThis paper proposes a new fault tolerant cache organization capable of dynamically mapping the in-use defective locations in a processor cache to an auxiliary parallel memory, creating a defect-free view of the cache for the processor. While voltage scaling has a super-linear effect on reducing power, it exponentially increases the defect rate in memory. The ability of the proposed cache organization to tolerate a large number of defects makes it a perfect candidate for voltage-scalable architectures, especially in smaller geometries where manufacturing induced process variation (MIPV) is expected to rapidly increase. The introduced fault tolerant architecture consumes little energy and area overhead, but enables the system to operate correctly and boosts the system performance close to a defect-free system. Power savings of over 40% is reported on standard benchmarks while the performance degradation is maintained below 1%. Avesta Sasan, Houman Homayoun, Ahmed M. Eltawil, Fadi J. Kurdahi |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2010 | E < MC2: less energy through multi-copy cacheabstractCaches are known to consume a large part of total microprocessor power. Traditionally, voltage scaling has been used to reduce both dynamic and leakage power in caches. However, aggressive voltage reduction causes process-variation-induced failures in cache SRAM arrays, which compromise cache reliability. We present Multi-Copy Cache (MC2), a new cache architecture that achieves significant reduction in energy consumption through aggressive voltage scaling, while maintaining high error resilience (reliability) by exploiting multiple copies of each data item in the cache. Unlike many previous approaches, MC2 does not require any error map characterization and therefore is responsive to changing operating conditions (e.g., Vdd-noise, temperature and leakage) of the cache. MC2 also incurs significantly lower overheads compared to other ECC-based caches. Our experimental results on embedded benchmarks demonstrate that MC2 achieves up to 60% reduction in energy and energy-delay product (EDP) with only 3.5% reduction in IPC and no appreciable area overhead. Arup Chakraborty, Houman Homayoun, Amin Khajeh, Nikil Dutt, Ahmed M. Eltawil, Fadi J. Kurdahi |
CASES | 6 |
| 2010 | Exploiting Architectural Similarities and Mode Sequencing in Joint Cost Optimization of Multi-mode FIR FiltersabstractWe present a new approach to designing multi-mode FIR filters in FPGAs based on partitioning a design into shared, mode-independent blocks and reconfigurable, mode-specific regions. We also provide a theoretical formulation for finding an optimal sequence of modes that minimizes a joint cost function related to area and reconfiguration overhead. Our results show that for a group of template matching filters, appropriate mode sequencing can reduce area by up to 15% and reconfiguration overhead by as much as 26%. Amir Hossein Gholamipour, Fadi J. Kurdahi, Ahmed M. Eltawil, Mazen A. R. Saghir |
FPL | 2 |
| 2010 | A Unified Hardware and Channel Noise Model for Communication SystemsabstractThis paper presents a single, scalable, unified statistical model that accurately reflects the impact of random embedded memory failures due to power management policies on the overall performance of a communication system. The proposed framework enables system designers to efficiently and accurately determine the effectiveness of novel power management techniques and algorithms that are designed to manage both hardware failure and communication channel noise, without the added cost of lengthy system simulations that are inherently limited and suffer from lack of scalability. Furthermore, the proposed framework facilitates performing both cross layer and intra layer tradeoffs where the faulty hardware can be treated as error-free hardware thus creating a much richer design space of power, performance and reliability. Amin Khajeh, Kiarash Amiri, Muhammed S. Khairy, Ahmed M. Eltawil, Fadi J. Kurdahi |
GLOBECOM | 5 |
| 2010 | RELOCATE: Register File Local Access Pattern Redistribution Mechanism for Power and Thermal Management in Out-of-Order Embedded Processor
Houman Homayoun, Aseem Gupta, Alexander V. Veidenbaum, Avesta Sasan, Fadi J. Kurdahi, Nikil Dutt |
HiPEAC | 5 |
| 2010 | Effect of body biasing on embedded SRAM failureabstractThis paper studies the tradeoffs when using body biasing as a power consumption modulator for Static Random Access Memories (SRAM) in term of reliability versus power consumption. We show that for fault tolerant applications such as wireless applications and multimedia, utilizing body biasing combined with voltage scaling can result in up to 47% power saving compared to the nominal case, while, for the same scenario, utilizing only voltage scaling will result in 20% power saving in the memory compared to nominal case. Amin Khajeh, Ahmed M. Eltawil, Fadi J. Kurdahi |
ISCAS | 3 |
| 2010 | Low-Power Multimedia System Design by Aggressive Voltage ScalingabstractMobile multimedia systems are growing in complexity and scalability and, correspondingly, in their implementation challenges. By design, these systems have built-in error resilience that has been exploited in many different compression and transmission schemes mainly as a quality tradeoff. This paper proposes a paradigm shift in utilizing error resilience in an application-aware method for reducing the power consumption of memories in such systems by aggressively scaling the supply voltage beyond what is currently considered as ¿safe¿ operating conditions while maintaining performance. Results on H.264 decoders show that power savings of more than 40% are possible. Fadi J. Kurdahi, Ahmed M. Eltawil, Kang Yi, Stanley Cheng, Amin Khajeh |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2010 | Evaluating Carbon Nanotube Global Interconnects for Chip Multiprocessor ApplicationsabstractIn ultra-deep submicrometer (UDSM) technologies, the current paradigm of using copper (Cu) interconnects for on-chip global communication is rapidly becoming a serious performance bottleneck. In this paper, we perform a system level evaluation of Carbon Nanotube (CNT) interconnect alternatives that may replace conventional Cu interconnects. Our analysis explores the impact of using CNT global interconnects on the performance and energy consumption of several multi-core chip multiprocessor (CMP) applications. Results from our analysis indicate that with improvements in fabrication technology, CNT-based global interconnects can significantly outperform Cu-based global interconnects. Sudeep Pasricha, Fadi J. Kurdahi, Nikil Dutt |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2010 | CAPPS: A Framework for Power-Performance Tradeoffs in Bus-Matrix-Based On-Chip Communication Architecture SynthesisabstractOn-chip communication architectures have a significant impact on the power consumption and performance of emerging chip multiprocessor (CMP) applications. However, customization of such architectures for an application requires the exploration of a large design space. Designers need tools to rapidly explore and evaluate relevant communication architecture configurations exhibiting diverse power and performance characteristics. In this paper, we present an automated framework for fast system-level, application-specific, power-performance tradeoffs in a bus matrix communication architecture synthesis (CAPPS). Our study makes two specific contributions. First, we develop energy models for system-level exploration of bus matrix communication architectures. Second, we incorporate these models into a bus matrix synthesis flow that enables designers to efficiently explore the power-performance design space of different bus matrix configurations. Experimental results show that our energy macromodels incur less than 5% average cycle energy error across 180-65 nm technology libraries. Our early system-level power estimation approach also shows a significant speedup ranging from 1000 to 2000× when compared with detailed gate-level power estimation. Furthermore, on applying our synthesis framework to three industrial networking CMP applications, a tradeoff space that exhibits up to 20% variation in power and up to 40% variation in performance is generated, demonstrating the usefulness of our approach. Sudeep Pasricha, Young-Hwan Park, Fadi J. Kurdahi, Nikil Dutt |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2009 | Dynamically reconfigurable on-chip communication architectures for multi use-case chip multiprocessor applicationsabstractThe phenomenon of digital convergence and increasing application complexity today is motivating the design of chip multiprocessor (CMP) applications with multiple use cases. Most traditional on-chip communication architecture design techniques perform synthesis and optimization only for a single use-case, which may lead to sub-optimal design decisions for multi-use case applications. In this paper we present a framework to generate a dynamically reconfigurable crossbar-based on-chip communication architecture that can support multiple use-case bandwidth and latency constraints. Our framework generates on-chip communication architectures with a low cost, low power dissipation, and with minimal reconfiguration overhead. Results of applying our framework on several networking CMP applications show that our approach is able to generate a crossbar solution with significantly lower cost (2.4times to 3.8times), and lower power dissipation (1.5times to 3.1times), compared to the best previously proposed approach. Sudeep Pasricha, Nikil Dutt, Fadi J. Kurdahi |
ASP-DAC | 3 |
| 2009 | A fault tolerant cache architecture for sub 500mV operation: resizable data composer cache (RDC-cache)abstractIn this paper we introduce Resizable Data Composer-Cache (RDC-Cache). This novel cache architecture operates correctly at sub 500 mV in 65 nm technology tolerating large number of Manufacturing Process Variation induced defects. Based on a smart relocation methodology, RDC-Cache decomposes the data that is targeted for a defective cache way and relocates one or few word to a new location avoiding a write to defective bits. Upon a read request, the requested data is recomposed through an inverse operation. For the purpose of fault tolerance at low voltages the cache size is reduced, however, in this architecture the final cache size is considerably higher compared to previously suggested resizable cache organizations [2][3]. The following three features a) compaction of relocated words, b)ability to use defective words for fault tolerance and c) "linking" (relocating the defective word to any row in the next bank), allows this architecture to achieve far larger fault tolerance in comparison to [2][3]. In high voltage mode, the fault tolerant mechanism of RDC-Cache is turned-off with minimal (0.91%) latency overhead compared to a traditional cache. Avesta Sasan, Houman Homayoun, Ahmed M. Eltawil, Fadi J. Kurdahi |
CASES | 4 |
| 2009 | TRAM: A tool for Temperature and Reliability Aware Memory DesignabstractMemories are increasingly dominating Systems on Chip (SoC) designs and thus contribute a large percentage of the total system's power dissipation, area and reliability. In this paper, we present a tool which captures the effects of supply voltage Vddand temperature on memory performance and their interrelationships. We propose a Temperature- and Reliability- Aware Memory Design (TRAM) approach which allows designers to examine the effects of frequency, supply voltage, power dissipation, and temperature on reliability in a mutually interrelated manner. Our experimental results indicate that thermal unaware estimation of probability of error can be off by at least two orders of magnitude and up to five orders of magnitude from the realistic, temperature-aware cases. We also observed that thermal aware Vddselection using TRAM can reduce the total power dissipation by up to 2.5times while attaining an identical predefined limit on errors. Amin Khajeh, Aseem Gupta, Nikil Dutt, Fadi J. Kurdahi, Ahmed M. Eltawil, Kamal S. Khouri, Magdy S. Abadir |
DATE | 4 |
| 2009 | Process Variation Aware SRAM/Cache for aggressive voltage-frequency scalingabstractThis paper proposes a novel Process Variation Aware SRAM architecture designed to inherently support voltage scaling. The peripheral circuitry of the SRAM is modified to selectively allow overdriving a wordline which contains weak cell(s). This architecture allows reducing the power on the entire array; however it selectively trades power for correctness when rows containing weak cells are accessed. The cell sizing is designed to assure successful read operations. This avoids flipping the content of the cells when the wordline is overdriven. Our simulations report 23% to 30% improvement in cell access time and 31% to 51% improvement in cell write time in overdriven wordlines. Total area overhead is negligible (4%). Low voltage operation achieves more than 40% reduction in dynamic power consumption and approximately 50% reduction in leakage power consumption. Avesta Sasan, Houman Homayoun, Ahmed M. Eltawil, Fadi J. Kurdahi |
DATE | 4 |
| 2009 | Size-Reconfiguration Delay Tradeoffs for a Class of DSP Blocks in Multi-mode Communication SystemsabstractIn this paper we propose a spectrum of designs for filters in multi-mode communication systems. The proposed designs lie in between generic filter to fully optimized coefficient specific filter. For each design a reconfigurable section and a static section are defined. We propose an algorithm to optimize the size of the reconfigurable section of each design independently. We also propose another algorithm that optimizes the reconfiguration time overhead for a given sequence of designs. The results of our experiments show the trade-off between area and reconfiguration delay in the design space. Amir Hossein Gholamipour, Hamid Eslami, Ahmed M. Eltawil, Fadi J. Kurdahi |
FCCM | 4 |
| 2009 | System-level PVT variation-aware power exploration of on-chip communication architecturesabstractWith the shift towards deep submicron (DSM) technologies, the increase in leakage power and the adoption of power-aware design methodologies have resulted in potentially significant variations in power consumption under different process, voltage, and temperature (PVT) corners. In this article, we first investigate the impact of PVT corners on power consumption at the system-on-chip (SoC) level, especially for the on-chip communication infrastructure. Given a target technology library, we then show how it is possible to “scale up” and abstract the PVT variability at the system level, allowing characterization of the PVT-aware design space early in the design flow. We conducted several experiments to estimate power for PVT corner cases, at the gate level, as well as at the higher system level. Our preliminary results are very interesting, and indicate that (i) there are significant variations in power consumption across PVT corners; and (ii) the PVT-aware power estimation problem may be amenable to a reasonably simple abstraction at the system level. Sudeep Pasricha, Young-Hwan Park, Nikil Dutt, Fadi J. Kurdahi |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2009 | A Low Power JPEG2000 Encoder With Iterative and Fault Tolerant Error ConcealmentabstractThis paper presents a novel approach to reduce power in multimedia devices. Specifically, we focus on JPEG2000 as a case study. This paper indicates that by utilizing the in-built error resiliency of multimedia content, and the disjoint nature of the encoding and decoding processes, ultra low power architectures that are hardware fault tolerant can be conceived. These architectures utilize aggressive voltage scaling to conserve power at the encoder side while incurring extra processing requirements at the decoder to blindly detect and correct for encoder hardware induced errors. Simulations indicate a reduction of up to 35% in encoder power depending on the choice of technology for a 65-nm CMOS process. Avesta Sasan, Amin Khajeh, Ahmed M. Eltawil, Fadi J. Kurdahi |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2007 | LEAF: A System Level Leakage-Aware Floorplanner for SoCsabstractProcess scaling and higher leakage power have resulted in increased power densities and elevated die temperatures. Due to the interdependence of temperature and leakage power, we observe that the floorplan has an impact on both the temperatures and the leakage of the IP-blocks in a system on chip (SoC). Hence, in this paper we propose a novel system level leakage aware floorplanner (LEAF) which optimizes floorplans for temperature-aware leakage power along with the traditional metrics of area and wire length. Our floorplanner takes a SoC netlist and the dynamic power profile of functional blocks to determine a placement while optimizing for temperature dependent leakage power, area, and wire length. To demonstrate the effectiveness of LEAF, we implemented our methodology on ten industrial SoC designs from Freescale Semiconductor Inc. and evaluated the trade-off between leakage power and area. We observed up to 190% difference in the leakage power between leakage-unaware and leakage aware floorplanning. Aseem Gupta, Nikil Dutt, Fadi J. Kurdahi, Kamal S. Khouri, Magdy S. Abadir |
ASP-DAC | 3 |
| 2007 | Exploiting Fault Tolerance Towards Power Efficient Wireless Multimedia ApplicationsabstractThis paper exploits the inherent redundancy available in wireless multimedia systems to tradeoff system redundancy versus power consumption. It is shown that aggressive power management techniques result in hardware failures that are mostly localized to embedded memories. By isolating these errors and utilizing system level fault tolerance techniques, we show that a reduction in power of up to 38% is possible in an H.264 decoder. This is coupled with a reduction of up to 20% in the underlying 3GPP WCDMA modem. With the proliferation of wideband wireless systems, an inevitable result is increased demand on receiving high quality, high bandwidth video. Based on these trends, one can identify three main challenges facing mobile, multimedia and communication designers. The first and foremost is power consumption which is on the rise due to the complex algorithms necessary to enable broadband multimedia wireless communication in dispersive and highly mobile environments. The second challenge is technology related, where scaling is both an enabler and a limiter. It enables unprecedented integration, including the ability to integrate large memories on chip, with the downside being a penalty in leakage power as well as reliability. Finally, the third challenge, cost is typically a major restricting factor. Designers are faced with the daunting dilemma of generating high yielding architectures that integrate vast amounts of logic and memories in a minimum die size with minimum power consumption. In this paper, we present an alternative approach to designing systems that have built in inherent redundancy. Amin Khajeh, Ahmed M. Eltawil, Fadi J. Kurdahi |
CCNC | 3 |
| 2007 | Error-Aware DesignabstractThe universal underlying assumption made today is that systems on chip must maintain 100% correctness regardless of the application. This work advocates the concept that some applications - by construction - are inherently error tolerant and therefore do not require this strict bound of 100% correctness. In such cases, it is possible to exploit this tolerance by aggressively reducing the supply voltage, thereby reducing power consumption significantly. This approach is demonstrated on several case studies in imaging, video and wireless communication fields. Fadi J. Kurdahi, Ahmed M. Eltawil, Amin Khajeh, Avesta Sasan, Stanley Cheng |
DSD | 1 |
| 2007 | Power Management for Cognitive Radio PlatformsabstractThis paper discusses how the cognitive radio concept can be extended to allow the system not only to manage shared resources such as spectrum, but to use this knowledge to optimize the overall system power consumption. We introduce a case study of video over wireless via a 3G WCDMA modem connected to an H.264 decoder. We show that by utilizing knowledge about the communication channel, a savings of more than 20% of the overall system power is possible while maintaining a required quality of service. Amin Khajeh, Shih-Yang Cheng, Ahmed M. Eltawil, Fadi J. Kurdahi |
GLOBECOM | 4 |
| 2007 | Limits on voltage scaling for caches utilizing fault tolerant techniquesabstractThis paper proposes a new low power cache architecture that utilizes fault tolerance to allow aggressively reduced voltage levels. The fault tolerant overhead circuits consume little energy, but enable the system to operate correctly and boost the system performance to close to defect free operation. Overall, power savings of over 40% are reported on standard benchmarks. Avesta Sasan, Amin Khajeh, Ahmed M. Eltawil, Fadi J. Kurdahi |
ICCD | 4 |
| 2007 | System level power estimation methodology with H.264 decoder prediction IP case studyabstractThis paper presents a methodology to generate a hierarchy of power models for power estimation of custom hardware IP blocks, enabling a trade-off between power estimation accuracy, modeling effort and estimation speed. Our power estimation approach enables several novel system-level explorations - such as observing the effect of clock gating, and the effects of tweaking application-level parameters on system power - with an estimation accuracy that is close to the gate-level. We implemented our methodology on an H.264 video decoder prediction IP case study, created power models, and evaluated the effects of varying design parameters (e.g., clock gating, IIP frame ratios, quantization), allowing rapid system-level power exploration of these design parameters. Young-Hwan Park, Sudeep Pasricha, Fadi J. Kurdahi, Nikil Dutt |
ICCD | 3 |
| 2007 | Fault Tolerant Approaches Targeting Ultra Low Power Communications System DesignabstractThis paper presents a new approach to co-designing communication systems and their respective hardware architectures. We show that by taking into account the specific needs and assumptions of the algorithms running on hardware, the circuit specifications targeted by ASIC engineers can be relaxed. This in turn leads to optimal designs in both power consumption and robustness. A case-study of a complete WCDMA modem incorporating this approach shows a savings of 23% in embedded memory power consumption and a total of 13% power savings for the whole system. Amin Khajeh, Ahmed M. Eltawil, Fadi J. Kurdahi |
VTC Spring | 3 |
| 2007 | A scalable embedded JPEG 2000 architecture
Fadi J. Kurdahi |
J. Syst. Archit. | 3 |
| 2006 | Compile-time area estimation for LUT-based FPGAsabstractThe Cameron Project has developed a system for compiling codes written in a high-level language called SA-C, to FPGA-based reconfigurable computing systems. In order to exploit the parallelism available on the FPGAs, the SA-C compiler performs a large number of optimizations such as full loop unrolling, loop fusion and strip-mining. However, since the area on an FPGA is limited, the compiler needs to know the effect of compiler optimizations on the FPGA area; this information is typically not available until after the synthesis and place and route stage, which can take hours. In this article, we present a compile-time area estimation technique to guide SA-C compiler optimizations. We demonstrate our technique for a variety of benchmarks written in SA-C. Experimental results show that our technique predicts the area required for a design to within 2.5% of actual for small image processing operators and to within 5.0% for larger benchmarks. The estimation time is in the order of milliseconds, compared with minutes for the synthesis tool. Dhananjay Kulkarni, Walid A. Najjar, Robert Rinker, Fadi J. Kurdahi |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2005 | On combining iteration space tiling with data space tiling for scratch-pad memory systemsabstractMost previous studies on tiling concentrate on iteration space only for cache-based memory systems. However, more and more real-time embedded systems are adopting Scratch-Pad Memories (SPMs) which emphasize on the management of data flow through data-oriented tiling. In this paper, we analyze the relationships between iteration space I and data space D, proposing a preliminary classification based on subscript functions. An important real-life application, matrix multiply, is selected to illustrate how we combine the mismatched iteration space tiling with data space tiling for optimal solutions. Fadi J. Kurdahi |
ASP-DAC | 2 |
| 2003 | Automatic compilation to a coarse-grained reconfigurable system-opn-chipabstractThe rapid growth of device densities on silicon has made it feasible to deploy reconfigurable hardware as a highly parallel computing platform. However, one of the obstacles to the wider acceptance of this technology is its programmability. The application needs to be programmed in hardware description languages or an assembly equivalent, whereas most application programmers are used to the algorithmic programming paradigm. SA-C has been proposed as an expression-oriented language designed to implicitly express data parallel operations. The Morphosys project proposes an SoC architecture consisting of reconfigurable hardware that supports a data-parallel, SIMD computational model. This paper describes a compiler framework to analyze SA-C programs, perform optimizations, and automatically map the application onto the Morphosys architecture. The mapping process is static and it involves operation scheduling, processor allocation and binding, and register allocation in the context of the Morphosys architecture. The compiler also handles issues concerning data streaming and caching in order to minimize data transfer overhead. We have compiled some important image-processing kernels, and the generated schedules reflect an average speedup in execution times of up to 6× compared to the execution on 800 MHz Pentium III machines. Girish Venkataramani, Walid A. Najjar, Fadi J. Kurdahi, Nader Bagherzadeh, A. P. Wim Böhm, Jeffrey Hammes |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2002 | A Complete Data Scheduler for Multi-Context Reconfigurable ArchitecturesabstractA new technique is presented in this paper to improve the efficiency of data scheduling for multi-context reconfigurable architectures targeting multimedia and DSP applications. The main goal is to improve the applications execution time minimizing external memory transfers. Some amount of on-chip data storage is assumed to be available in the reconfigurable architecture. Therefore the Complete Data Scheduler tries to optimally exploit this storage, saving data and result transfers between on-chip and external memories. In order to do this, specific algorithms for data placement and replacement have been designed. We also show that a suitable data scheduling could decrease the number of transfers required to implement the dynamic reconfiguration of the system. Marcos Sánchez-Élez Martín, Milagros Fernández, Rafael Maestre-Ferriz, Román Hermida, Nader Bagherzadeh, Fadi J. Kurdahi |
DATE | 6 |
| 2002 | MorphoSys: A Coarse Grain Reconfigurable Architecture for Multimedia Applications (Research Note)
Hooman Parizi, Afshin Niktash, Nader Bagherzadeh, Fadi J. Kurdahi |
Euro-Par | 4 |
| 2002 | Fast Area Estimation to Support Compiler Optimizations in FPGA-Based Reconfigurable SystemsabstractSeveral projects have developed compiler tools that translate high-level languages down to hardware description languages for mapping onto FPGA-based reconfigurable computers. These compiler tools can apply extensive transformations that exploit the parallelism inherent in the computations. However, the transformations can have a major impact on the chip area (number of logic blocks) used on the FPGA. It is imperative therefore that the compiler user be provided with feedback indicating how much space is being used. In this paper we present a fast compile-time area estimation technique to guide the compiler optimizations. Experimental results show that our technique achieves an accuracy within 2.5% for small image-processing operators, and within 5.0% for larger benchmarks, as compared to the usual post-compilation synthesis tool estimations. The estimation time is in the order of milliseconds as compared to several minutes for a synthesis tool. Dhananjay Kulkarni, Walid A. Najjar, Robert Rinker, Fadi J. Kurdahi |
FCCM | 4 |
| 2002 | Guest editorial special issue on system synthesis
Fadi J. Kurdahi |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2001 | A compiler framework for mapping applications to a coarse-grained reconfigurable computer architectureabstractThe rapid growth of silicon densities has made it feasible to deploy reconfigurable hardware as a highly parallel computing platform. However, in most cases, the application needs to be programmed in hardware description or assembly languages, whereas most application programmers are familiar with the algorithmic programming paradigm. SA-C has been proposed as an expression-oriented language designed to implicitly express data parallel operations. Morphosys is a reconfigurable system-on-chip architecture that supports a data-parallel, SIMD computational model. This paper describes a compiler framework to analyze SA-C programs, perform optimizations, and map the application onto the Morphosys architecture. The mapping process involves operation scheduling, resource allocation and binding and register allocation in the context of the Morphosys architecture. The execution times of some compiled image-processing kernels can achieve up to 42x speed-up over an 800 MHz Pentium III machine. Girish Venkataramani, Walid A. Najjar, Fadi J. Kurdahi, Nader Bagherzadeh, A. P. Wim Böhm |
CASES | 3 |
| 2001 | Power-Aware Scheduling under Timing Constraints for Mission-Critical Embedded SystemsabstractPower-aware systems are those that must make the best use of available power. They subsume traditional low-power systems in that they must not only minimize power when the budget is low, but also deliver high performance when required. This paper presents a new scheduling technique for supporting the design and evaluation of a class of power-aware systems in mission-critical applications. It satisfies stringent min/max timing and max power constraints. It also makes the best effort to satisfy the min power constraint in an attempt to fully utilize free power or to control power jitter. Experimental results show that our scheduler can improve performance and reduce energy cost simultaneously compared to hand-crafted designs for previous missions. This tool forms the basis of the IM-PACCT system-level framework that will enable designers to explore many power-performance trade-offs with confidence. 1. Jinfeng Liu 0006, Pai H. Chou, Nader Bagherzadeh, Fadi J. Kurdahi |
DAC | 4 |
| 2001 | Kernel scheduling techniques for efficient solution space exploration in reconfigurable computing
Rafael Maestre-Ferriz, Fadi J. Kurdahi, Milagros Fernández, Román Hermida, Nader Bagherzadeh, Hartej Singh |
J. Syst. Archit. | 2 |
| 2001 | A framework for reconfigurable computing: task scheduling and context managementabstractDynamically reconfigurable architectures are emerging as a viable design alternative to implement a wide range of computationally intensive applications. At the same time, an urgent necessity has arisen for support tool development to automate the design process and achieve optimal exploitation of the architectural features of the system. Task scheduling and context (configuration) management become very critical issues in achieving the high performance that digital signal processing (DSP) and multimedia applications demand. This article proposes a strategy to automate the design process which considers all possible optimizations that can be carried out at compilation time, regarding context and data transfers. This strategy is general in nature and could be applied to different reconfigurable systems. We also discuss the key aspects of the scheduling problem in a reconfigurable architecture such as MorphoSys. In particular, we focus on a task scheduling methodology for DSP and multimedia applications, as well as the context management and scheduling optimizations. Rafael Maestre-Ferriz, Fadi J. Kurdahi, Milagros Fernández, Román Hermida, Nader Bagherzadeh, Hartej Singh |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2000 | MorphoSys: case study of a reconfigurable computing system targeting multimedia applicationsabstractIn this paper, we present a case study for the design, programming and usage of a reconfigurable system-on-chip, MorphoSys, which is targeted at computation-intensive applications. This 2-million transistor design combines a reconfigurable array of cells with a RISC processor core and a high bandwidth memory interface. The system architecture, software tools including a scheduler for reconfigurable systems, and performance analysis (with impressive speedups) for target applications are described. Hartej Singh, Eliseu M. Chaves Filho, Rafael Maestre-Ferriz, Ming-Hau Lee, Fadi J. Kurdahi, Nader Bagherzadeh |
DAC | 6 |
| 2000 | Optimal vs. Heuristic Approaches to Context Scheduling for Multi-Context Reconfigurable ArchitecturesabstractThis paper describes a methodology to efficiently obtain a solution to the problem of context scheduling for multi-context reconfigurable architectures, regarding the minimization of context loading overhead. The target applications are assumed to be periodic, since it is a typical feature of many DSP and multimedia applications. This paper considers the trade-off between achievable system performance and algorithm efficiency. It has been developed as a part of an automated design environment for reconfigurable systems. Rafael Maestre-Ferriz, Milagros Fernández, Román Hermida, Fadi J. Kurdahi, Nader Bagherzadeh, Hartej Singh |
FCCM | 4 |
| 2000 | Optimal vs. Heuristic Approaches to Context Scheduling for Multi-Context Reconfigurable ArchitecturesabstractThis paper describes a methodology to efficiently obtain a solution to the problem of context scheduling for multi-context reconfigurable architectures, regarding the minimization of context loading overhead. The target applications are assumed to be periodic, since it is a typical feature of many DSP and multimedia applications. This work considers the trade-off between achievable system performance and algorithm efficiency. It has been developed as a part of an automated design environment for reconfigurable systems. Rafael Maestre-Ferriz, Milagros Fernández, Román Hermida, Fadi J. Kurdahi, Nader Bagherzadeh, Hartej Singh |
ICCD | 4 |
| 2000 | MorphoSys: An Integrated Reconfigurable System for Data-Parallel and Computation-Intensive ApplicationsabstractThis paper introduces MorphoSys, a reconfigurable computing system developed to investigate the effectiveness of combining reconfigurable hardware with general-purpose processors for word-level, computation-intensive applications. MorphoSys is a coarse-grain, integrated, and reconfigurable system-on-chip, targeted at high-throughput and data-parallel applications. It is comprised of a reconfigurable array of processing cells, a modified RISC processor core, and an efficient memory interface unit. This paper describes the MorphoSys architecture, including the reconfigurable processor array, the control processor, and data and configuration memories. The suitability of MorphoSys for the target application domain is then illustrated with examples such as video compression, data encryption and target recognition. Performance evaluation of these applications indicates improvements of up to an order of magnitude (or more) on MorphoSys, in comparison with other systems. Hartej Singh, Ming-Hau Lee, Fadi J. Kurdahi, Nader Bagherzadeh, Eliseu M. Chaves Filho |
IEEE Trans. Computers | 4 |
| 1999 | Kernel Scheduling in Reconfigurable ComputingabstractReconfigurable computing is a flexible way of facing with a single device a wide range of applications with a good level of performance. This area of computing involves different issues and concepts when compared with conventional computing systems. One of these concepts is context lending. The context refers to the coded configuration information to implement a particular circuit behaviour. An important problem for reconfigurable computing is the scheduling of a group of kernels (sub-tasks) that constitute a complex application for minimum execution time. In this paper, we show how the different execution orders for these sub-tasks may result in varying levels of performance. We formulate an analytical approach and present a solution for this new problem through this work. Rafael Maestre-Ferriz, Fadi J. Kurdahi, Nader Bagherzadeh, Hartej Singh, Román Hermida, Milagros Fernández |
DATE | 2 |
| 1999 | The MorphoSys Parallel Reconfigurable System
Hartej Singh, Ming-Hau Lee, Nader Bagherzadeh, Fadi J. Kurdahi, Eliseu M. Chaves Filho |
Euro-Par | 5 |
| 1999 | High-level synthesis of recoverable VLSI microarchitecturesabstractTwo algorithms that combine the operations of scheduling and recovery-point insertion for high-level synthesis of recoverable microarchitectures are presented. The first uses a prioritized cost function in which functional unit (FU) cost is minimized first and register cost second. The second algorithm minimizes a weighted sum of FU and register costs. Both algorithms are optimal according to their respective cost functions and require less than 10 min of central processing unit (CPU) time on widely used high-level synthesis benchmarks. The best previous result reported several hours of CPU time for some of the same benchmarks on a computer of similar computational power. Douglas M. Blough, Fadi J. Kurdahi, Seong Yong Ohm |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1999 | Accurate prediction of quality metrics for logic level designs targeted toward lookup-table-based FPGAsabstractThe importance of efficient area and timing estimation techniques is well-established in high-level synthesis (HLS) since it allows more efficient exploration of the design space while providing HLS tools with the capability of predicting the effects of technology-specific tools on the design space. Much of the previous work has focused on estimation techniques that use very simple cost models based solely on the gate and/or literal count. Those models are not accurate enough to allow effective design space exploration since the effects of interconnect can indeed dominate the final design cost. The situation becomes even worse when the design is targeted to field-programmable gate array (FPGA) technologies since the wire delay may contribute up to 60% of the overall design delay. In this paper, we present an approach of estimating area and timing for lookup-table-based FPGAs that takes into account not only gate area and delay, but also the wiring effects. We select the Xilinx XC4000 series as our main concentration because of their popularity. We tested our estimator with several benchmarks and the results show that we receive accurate area and timing estimates efficiently. Fadi J. Kurdahi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1998 | Layout-Driven High Level Synthesis for FPGA Based ArchitecturesabstractIn this paper, we address the problem of layout-driven scheduling-binding as these steps have a direct relevance on the final performance of the design. The importance of effective and efficient accounting of layout effects is well-established in High-Level Synthesis (HLS), since it allows more efficient exploration of the design space and the generation of solutions with predictable metrics. This feature is highly desirable in order to avoid unnecessary iterations through the design process. By producing not only an RTL netlist but also an approximate physical topology of implementation at the chip level, we ensure that the solution will perform at the predicted metric once implemented, thus avoiding unnecessary delays in the design process. Fadi J. Kurdahi |
DATE | 2 |
| 1998 | On the Characterization of Multi-Point Nets in Electronic DesignsabstractImportant layout properties of electronic designs include interconnection length values, clock speed, area requirements, and power dissipation. A reliable estimation of those properties is essential for improving placement and routing techniques for digital circuits. Previous work on estimating design properties failed to take multi-point nets into account. All nets were assumed to be 2-point nets (especially for estimating the number of nets). In this paper we aim at characterizing multi-point nets in electronic designs. We develop a model for the behaviour of multi-point nets during the partitioning process. The resulting distribution of nets over their net degree is validated through comparison with benchmark data. Dirk Stroobandt, Fadi J. Kurdahi |
Great Lakes Symposium on VLSI | 2 |
| 1997 | ChipEst-FPGA: a tool for chip level area and timing estimation of lookup table based FPGAs for high level applicationsabstractThe importance of efficient area and timing estimation techniques for hierarchical design methodology is well-established in High-Level Synthesis (HLS), since the estimation allows more realistic exploration of the design space, and hierarchical design methodology matches well with HLS paradigm. In this paper, we present ChipEst-FPGA, a chip level estimator for designs implemented using a hierarchical design methodology for Lookup Table Based FPGAs. In FPGAs, the wire delay may contribute to a significant portion of the overall design delay. ChipEst-FPGA uses a realistic model which takes the component area/delay as well as wiring effects into account. We tested our ChipEst-FPGA on several benchmarks and the results show that we can get accurate area and timing estimates efficiently. Fadi J. Kurdahi |
ASP-DAC | 2 |
| 1997 | Optimal algorithms for recovery point insertion in recoverable microarchitecturesabstractThis paper considers the problem of automatic insertion of recovery points in recoverable microarchitectures. Previous work on this problem provided heuristic nonoptimal algorithms that attempted either to minimize computation time with a bounded hardware overhead or to minimize hardware overhead with a bounded computation time. In this paper, we present polynomial-time algorithms that provide provably optimal solutions for both of these formulations of the problem. These algorithms take as their input a scheduled control-data flow graph describing the behavior of the system, and they output either a minimum time or a minimum cost set of recovery point locations. We demonstrate the performance of our algorithms using some well-known benchmark control-data flow graphs. Over all parameter values for each of these benchmarks, our optimal algorithms are shown to perform as well as, and in many cases better than, the previously proposed heuristics. Douglas M. Blough, Fadi J. Kurdahi, Seong Yong Ohm |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1997 | A unified lower bound estimation technique for high-level synthesisabstractThe importance of effective lower bound estimation (LBE) techniques is well established in high-level synthesis (HLS) since it allows more efficient exploration of the design space while providing other HLS tools with the capability of predicting the effect of specific tools on the design space. Much of the previous work has focused on LBE techniques that use very simple cost models which primarily focus on the functional unit resources. With the push toward submicron technologies, simple models that use functional unit resources alone are not accurate enough to allow effective design space exploration since the effects of storage and interconnect can indeed dominate the cost function. In this paper, we present an integrated approach aimed at predicting lower bounds on hardware resources needed to implement a behavioral description within a given amount of time. Our area cost model accounts for storage (register) and interconnect resources (buses) in addition to functional resources. Our timing model uses a finer granularity that permits the modeling of functional unit, register, and interconnect delays. Our approach is integrated because we consider the dependencies between the different types of resources as well as the ordering in which the resources are allocated. We tested our technique for functional unit, storage, and interconnect requirements on several high-level synthesis benchmarks, and observed near-optimal results. We believe that our comprehensive LBE approach can lead to better quality HLS solutions in less time, and we demonstrate this approach in our paper. Seong Yong Ohm, Fadi J. Kurdahi, Nikil Dutt |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1997 | Layout-driven RTL binding techniques for high-level synthesis using accurate estimatorsabstractThe importance of effective and efficient accounting of layout effects is well established in High-Level Synthesis (HLS), since it allows more realistic exploration of the design space and the generation of solutions with predictable metrics. This feature is highly desirable in order to avoid unnecessary iterations through the design process. In this article, we address the problem of layout-driven register-transfer-level (RTL) binding as this step has a direct relevance to the final performance of the design. By producing not only an RTL design but also an approximate physical topology of the chip-level implementation, we ensure that the solution will perform at the predicted metric once implemented, thus avoiding unnecessary delays in the design process. Fadi J. Kurdahi |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 1994 | A performance driven logic synthesis system using delay estimatorabstractIn this paper, we develop a logic synthesis approach which relies on accurate design evaluation program to estimate the final design attributes such as layout speed. Given a candidate design implementation, an evaluation program is called upon to provide quick and accurate estimates of the critical path delay. This information is then used as a feedback to the logic optimization system. Based on this feedback, the system will "re-orient" itself toward a new direction for optimization. Such a scheme represents a more realistic way of generating optimal layout implementations.> Wei Kang Tsai, Fadi J. Kurdahi, Tzong-Dar Her, Champaka Ramachandran |
Great Lakes Symposium on VLSI | 3 |
| 1994 | Comprehensive lower bound estimation from behavioral descriptionsabstractIn this paper, we present a comprehensive technique for lower bound estimation (LBE) of resources from behavioral descriptions. Previous work has focused on LBE techniques that use very simple cost models which primarily focus on the functional unit resources. Our cost model accounts for storage resources in additionto functionalresources. Our timing model uses a finer granularity that permits the modeling of functional unit, register and interconnect delays. We tested our LBE technique for both functional unit and storage requirements on several high-level synthesis benchmarks and observed near-optimal results. 1 Seong Yong Ohm, Fadi J. Kurdahi, Nikil Dutt |
ICCAD | 2 |
| 1994 | On the intrinsic Rent parameter and spectra-based partitioning methodologiesabstractThe complexity of circuit designs has necessitated a top-down approach to layout synthesis. A large body of work shows that a good layout hierarchy, or partitioning tree, as measured by the associated Rent parameter, will correspond to an area-efficient layout. We define the intrinsic Rent parameter of a netlist to be the minimum possible Rent parameter of any partitioning tree for the netlist. Experimental results show that spectra-based ratio cut partitioning algorithms yield partitioning trees with the lowest observed Rent parameter over all benchmarks and over all algorithms tested. For examples where the intrinsic Rent parameter is known, spectral ratio cut partitioning yields a partitioning tree with Rent parameter essentially identical to this theoretical optimum. These results have deep implications with respect to both the choice of partitioning algorithms for top-down layout, as well as new approaches to layout area estimation. The paper concludes with directions for future research, including several promising techniques for fast estimation of the (intrinsic) Rent parameter.> Lars W. Hagen, Andrew B. Kahng, Fadi J. Kurdahi, Champaka Ramachandran |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 1994 | Combined topological and functionality-based delay estimation using a layout-driven approach for high-level applicationsabstractWe discuss the problem of accurate delay estimation of cell-based designs, prior to any physical design tasks. For this purpose, we require accurate wire-length estimates, since wire delays contribute significantly to the overall delay. We present a new technique for wire-length estimation based on a combination of analytical and constructive approaches. Given these wire-length estimates and the cell delays, it is possible to provide worst case delay paths in the design based on the circuit topology. We have also extended our technique to consider false paths, which provides a more accurate functionality based estimate that takes into account the estimated layout information. We validate our technique using the standard MCNC benchmarks. Our results indicate an average 7% accuracy in the worst case delay predictions for designs with up to about 2800 cells.> Champaka Ramachandran, Fadi J. Kurdahi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1993 | A logic synthesis system based on global dynamic extraction and flexible costabstractAn efficient algorithm for a logic synthesis system based on global dynamic extraction and flexible cost (GDEF) is described. The GDEF is designed to find the best common subexpression and to update the value of the other common subexpressions dynamically. The GDEF feedbacks approximate layout area obtained through an area-estimation program to guide the logic synthesis process. In this approach, the cost function is dependent on both literal count and estimated area.> Wei Kang Tsai, Fadi J. Kurdahi |
Great Lakes Symposium on VLSI | 3 |
| 1993 | On clustering for maximal regularity extractionabstractThe authors point out that proper usage of regularity in digital systems leads to efficient as well as economical designs. This important question of regularity extraction is examined, and a general and efficient methodology for component clustering based on the concept of structural regularity is presented. While the concept of regularity can be employed to simplify many problems in the area of design automation, system- and logic-level applications are emphasized here. The authors show how identifying clusters in a circuit can simplify two important CAD problems-system-level clustering and module (layout) generation. A prototype system based on these ideas has been built, and some real-life examples are considered for testing. The results are encouraging; they demonstrate the essential role such a system could play in aiding the high-level system designer. Research is under way to explore some of the other promising applications that such a system could have.> D. Sreenivasa Rao, Fadi J. Kurdahi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1993 | Evaluating layout area tradeoffs for high level applicationsabstractThe authors address the problem of evaluating area tradeoffs for VLSI layouts from high-level specifications (typically register-transfer level). An area prediction approach based on two models, analytical and constructive, are presented. A circuit design is partitioned recursively down to a level specified by the user, thus generating a slicing tree. An analytical model is then used to predict the shape function of each of the leaf subcircuits. By traversing the tree in post-order, the shape function of the entire layout design can be predicted constructively. This approach also permits the user to trade off the accuracy of the prediction versus the runtime of the predictor. Such a scheme is quite useful for high level design tasks. The authors show experimentally that the estimates obtained using this model are within 5% of the actual layout area for designs ranging from 125 to 12000 cells.> Fadi J. Kurdahi, Champaka Ramachandran |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1993 | Hierarchical design space exploration for a class of digital systemsabstractThis paper presents an architectural synthesis approach for a widely used class of digital systems characterized by inherent regularity in their description. This approach relies on a novel modeling or abstraction of the problem domain to facilitate a hierarchical solution method. The modeling is based on exploiting the inherent regularity in the system description to cluster its behavioral operations. The method emphasizes prudent postponement of design decisions until enough physical design information is available to estimate layout effects like wiring; we use well-known area-delay estimators for this purpose. The approach has the advantage that it keeps track of a set of potentially good candidate solutions, rather than narrowing down to a single solution very early in the design process. Through an extensive set of experiments on well-known DSP design examples, we demonstrate the advantages that such distinctive features have to offer; the impact of hierarchy on several important issues, such as interconnection area, extent of design space explored, etc., is presented.> D. Sreenivasa Rao, Fadi J. Kurdahi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1992 | Partitioning by Regularity Extraction
D. Sreenivasa Rao, Fadi J. Kurdahi |
DAC | 2 |
| 1992 | Accurate layout area and delay modeling for system level designabstractThe problem of estimating design quality measures to accurately reflect design tradeoffs and efficiently explore the design space is discussed. Specifically, interest is centered on predicting the layout area and delay of a given structural RT level design. Clearly, current RT level cost measures are highly simplified and do not reflect the real physical design. In order to establish a more realistic assessment of layout effects, a layout model which accurately and efficiently accounts for the effects of wiring and floorplanning on the area and performance layout of RT level designs is proposed. Benchmarking has shown that this model is quite accurate.> Champaka Ramachandran, Fadi J. Kurdahi, Daniel Gajski, Allen C.-H. Wu, Viraphol Chaiyakul |
ICCAD | 2 |
| 1991 | Automatic Synthesis of Time-Stationary Controllers for Pipelined Data PathsabstractThe authors present an approach for automatically synthesizing a time-stationary control scheme for a given pipelined data path. They have developed an efficient method of producing a control specification for the data path. A highly optimized FSM (finite state machine) controller implementation is obtained by partitioning so as to minimize the total controller area. The FSM controller is implemented using either PLAs or standard cells. The present approach has been compared to published work on FSM generation and optimization, and the results indicate large savings in total controller area.> James J. Kim, Fadi J. Kurdahi, Nohbyung Park |
ICCAD | 2 |
| 1989 | Module assignment and interconnect sharing in register-transfer synthesis of pipelined data pathsabstractThe authors present a novel approach to the problem of register-transfer (RT) design optimization of pipelined data paths. They perform module assignment with the goal of maximizing the interconnect sharing between RT-level components. The interconnect sharing task is modeled as a constrained clique partitioning problem. They have developed a fast and efficient polynomial time heuristic procedure to solve this problem. This procedure is 30-50 times faster than other existing heuristics while still producing better results for the authors' purposes.> Nohbyung Park, Fadi J. Kurdahi |
ICCAD | 2 |
| 1989 | Techniques for area estimation of VLSI layoutsabstractThe standard cell design style is investigated. Two probabilistic models are presented. The first model estimates the wiring space requirements in the routing channels between the cell rows. The second model estimates the number of feedthroughs that must be inserted in the cell rows to interconnect cells placed several rows apart. These models were implemented in the standard cell area estimation program PLEST (PLotting ESTimator). PLEST was used to estimate the areas of a set of 12 standard cell chips. In all cases, the estimates were accurate to within 10% of the actual areas. PLEST's estimation of a chip layout area takes only a few seconds to produce, as compared with more than 10 h to generate the chip layout itself using an industrial layout system.> Fadi J. Kurdahi, Alice C. Parker |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1987 | REAL: a program for REgister ALlocationabstractThis paper describes the REAL REgister ALlocation program. REAL uses a track assignment algorithm taken from channel routing called the Left Edge algorithm. REAL is optimal for non-pipelined designs with no conditional branches. It is thought that REAL is also optimal for designs with conditional branches, pipelined or not. Experimental results are included in the report, which illustrate the optimal solutions found by REAL. REAL is part of the ADAM Advanced Design AutoMation system, and will be used to process designs output from MAHA and Sehwa. Fadi J. Kurdahi, Alice C. Parker |
DAC | 1 |
| 1986 | PLEST: a program for area estimation of VLSI integrated circuitsabstractThis paper describes PLEST, a program for estimating the area of standard cell layouts as part of the more general ARREST area estimator embedded in the ADAM system. PLEST is based on a probabilistic model for placement of logic. Given various design parameters, PLEST generates a range of estimates for the possible shapes of the block layout. The program was applied to a set of six layouts. The estimated chip area is, for all six chips, within 10% of the measured area. Further research will be aimed at estimating layout area consumption starting from the register-transfer level design description. Fadi J. Kurdahi, Alice C. Parker |
DAC | 1 |
| 1984 | A general methodology for synthesis and verification of register-transfer designs
Alice C. Parker, Fadi J. Kurdahi, Mitch J. Mlinar |
DAC | 2 |