Shawki Areibi

dblp:56/5390 · also Shawki M. Areibi · DBLP profile ↗
← Back
39ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0003-4832-0911ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 21 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Design Space Exploration of Edge AI for Assistive Vision on Low-Cost Embedded Systems
N. Prasad, D. Patel, Gary William Grewal, Shawki Areibi
WoWMoM4
2025 Dual Graph Neural Networks for Optimizing Circuit Partitioning: A Synergistic Approach
abstract
Circuit partitioning is an NP-hard optimization problem that appears frequently throughout the Very Large Scale Integration (VLSI) design flow. This paper aims to investigate the performance of Graph Neural Networks (GNNs) in addressing this key problem. The proposed solution involves using two sequentially applied GNNs. The first GNN clusters the circuit’s graph representation to identify node features, which are then used as node embeddings by the second GNN to effectively partition the circuit. To evaluate the effectiveness of this approach, a direct performance comparison is conducted against the widely-used Sanchis multi-way partitioning heuristic. The empirical results are promising, demonstrating that GNNs achieve significant improvements in both quality of the partitioning and CPU runtime, especially for larger circuits and larger numbers of partitions.
A. Soroush, Shawki Areibi, Gary William Grewal, B. Yip
IJCNN2
2024 A High-Performance Routing Engine for Large-Scale FPGAs
abstract
Routing is the most time-consuming stage in the Field Programmable Gate Array (FPGA) design workflow. We propose a parallel routing technology, based on the Pathfinder algorithm, that enhances parallelism by dividing the search into two phases: one that tolerates overlaps and one that does not. Additional performance optimizations include an improved cost schedule, pruning the routing-resource graph, and selecting efficient data structures for modern CPUs. Evaluated using both the 2023 MLCAD and 2024 FPGA Routing Contest benchmarks, our router achieves average speedups of $6.2 \times$ and $5.2 \times$ compared to RWRoute and Vivado 2023.2, respectively.
Timothy Martin, Dani Maarouf, Gary William Grewal, Shawki Areibi
FPL4
2024 Invited Paper-Circuit Partitioning with Reinforcement Learning and Edge-Based Initialization
abstract
The Fiduccia-Mattheyses-Sanchis (FMS) algorithm is a widely used local search method for K-way circuit partitioning, but it’s prone to getting stuck in local minima. Traditionally, this has been addressed by running FMS multiple times with different random initial solutions, hoping for a better result. Building on our previous work with an RL-based local search method that helps FMS avoid these traps, this research explores a new approach: using constructive methods to generate superior initial solutions. We explored two such methods: NDE (node growing algorithm), a commonly used node-based method that maximizes node absorption, and NET (net growing algorithm), an edge-based approach that maximizes net absorption. By integrating NDE and NET with our RL-based local search, we’ve achieved significant improvements. Experiments on ISPD98/IBM benchmarks demonstrate that an edge-based approach provides higher-quality solutions for larger circuits and larger numbers of partitions. Combining these initial solutions with our RL-based approach further reduces the cutsize generate by the RL-based approach by up to 79.5%.
Ka Chuen Cheng, Umair F. Siddiqi, Gary William Grewal, Shawki Areibi
RSP4
2023 A Deterministic Parallel Routing Approach for Accelerating Pathfinder-based Algorithms
abstract
Routing is a time-consuming task in the FPGA design flow, and its task is to build non-overlapping routing trees for all nets. PathFinder is a popular routing algorithm, and it is implemented in the versatile-place-and-route (VPR) tool. The latest version of PathFinder, implemented in VPR 8.0, employs incremental routing in which it rip-up and re-route (RnR) only those branches of the routing trees that have a congested node or their delay has degraded significantly in the last iterations. The initial iterations have a very high workload (i.e., the number of branches to route), and the later ones have fewer branches to build. We propose a parallel-sequential hybrid router for PathFinder with incremental routing that applies deterministic parallel routing to a window of initial iterations having a high routing workload and sequential routing to the remaining iterations. It also uses an intelligent approach to select nets for sequential and parallel routing. Experiments conducted using Titan benchmarks show that it can improve the runtime of PathFinder by upto 32% with no significant degradation in solution quality.
Umair F. Siddiqi, Gary William Grewal, Shawki Areibi
VLSI-SoC3
2022 Guiding FPGA Detailed Placement via Reinforcement Learning
abstract
Detailed Placement (DP) is an important, but time-consuming, optimization step within the Field Programmable Gate Array (FPGA) design flow. Given a global placement, DP seeks to refine the global placement to improve the success of the subsequent routing step. In this paper, we show how Reinforcement Learning (RL) can be used to significantly reduce DP runtimes while maintaining Quality-of-Result (QoR). We develop 3 different RL models based on Tabular Q-Learning, Deep Q-Learning, and Actor-Critic. These models are evaluated by integrating them into GPlace3.0 – a state-of-the-art analytic FPGA placement tool – and tested using the 12 ISPD contest benchmarks. Our results show the models achieve total runtime improvements between 2x to 3.5x and similar QoR compared to GPlace3.0’s algorithmic-based detailed placer.
P. Esmaeili, Timothy Martin, Shawki Areibi, Gary William Grewal
VLSI-SoC3
2021 A Deep Learning Framework to Predict Routability for FPGA Circuit Placement
abstract
The ability to accurately and efficiently estimate the routability of a circuit based on its placement is one of the most challenging and difficult tasks in the Field Programmable Gate Array (FPGA) flow. In this article, we present a novel, deep learning framework based on a Convolutional Neural Network (CNN) model for predicting the routability of a placement. Since the performance of the CNN model is strongly dependent on the hyper-parameters selected for the model, we perform an exhaustive parameter tuning that significantly improves the model’s performance and we also avoid overfitting the model. We also incorporate the deep learning model into a state-of-the-art placement tool and show how the model can be used to (1) avoid costly, but futile, place-and-route iterations, and (2) improve the placer’s ability to produce routable placements for hard-to-route circuits using feedback based on routability estimates generated by the proposed model. The model is trained and evaluated using over 26K placement images derived from 372 benchmarks supplied by Xilinx Inc. We also explore several opportunities to further improve the reliability of the predictions made by the proposed DLRoute technique by splitting the model into two separate deep learning models for (a) global and (b) detailed placement during the optimization process. Experimental results show that the proposed framework achieves a routability prediction accuracy of 97% while exhibiting runtimes of only a few milliseconds.
Abeer Alhyari, Hannah Szentimrey, Ahmed Elshamli, Timothy Martin, Gary William Grewal, Shawki Areibi
ACM Trans. Reconfigurable Technol. Syst.6
2020 A Deep-Learning Framework for Predicting Congestion During FPGA Placement
abstract
The ability to quickly and accurately predict congestion has emerged as one of the most critical problems during placement. In this paper, we present DLCong, a deep learning congestion-estimation framework based on a convolutional encoder-decoder. Experimental results show that compared to MLCong, a state-of-the-art machine-learning based congestion-estimation model, DLCong achieves an almost 9% improvement in congestion accuracy, while exhibiting inference times of a few milliseconds. Moreover, the accuracy of DLCong scales better with increasing congestion compared to MLCong.
Dani Maarouf, Ahmed Elshamli, Timothy Martin, Gary William Grewal, Shawki Areibi
FPL5
2020 Multisource Domain Adaptation for Remote Sensing Using Deep Neural Networks
abstract
In applying machine learning to remote sensing problems, it is often the case that multiple training data sources, known as domains, are available for the same task. It is sample-inefficient to train separate models per domain, which motivates learning a single model from multiple sources. For example, the local climate zone (LCZ) classification problem that aims to produce per-pixel classifications of surface structure from remotely sensed images of urban and rural environments. These classification maps need to be generated for different cities at different times. To do this efficiently, available training data from different sources (i.e., cities) must be adapted for the task at hand. However, multisource domain adaptation (MDA) is a challenging problem and is particularly apparent when there are significant changes in the data distribution among these sources. In this article, we propose a scalable yet simple adaptive MDA (AMDA) framework to address this problem. AMDA is also capable of dealing with imbalanced data distributions among the sources more effectively than existing baselines. We also extend two techniques originally proposed for domain expansion (DE) to the task of DA. AMDA and the extended DE techniques are implemented and evaluated on the LCZ classification problem. Despite its simplicity, AMDA is able to achieve more than 12% improvement over the baseline.
Ahmed Elshamli, Graham W. Taylor, Shawki Areibi
IEEE Trans. Geosci. Remote. Sens.3
2020 Machine Learning for Congestion Management and Routability Prediction within FPGA Placement
abstract
Placement for Field Programmable Gate Arrays (FPGAs) is one of the most important but time-consuming steps for achieving design closure. This article proposes the integration of three unique machine learning models into the state-of-the-art analytic placement tool GPlace3.0 with the aim of significantly reducing placement runtimes. The first model, MLCong, is based on linear regression and replaces the computationally expensive global router currently used in GPlace3.0 to estimate switch-level congestion. The second model, DLManage, is a convolutional encoder-decoder that uses heat maps based on the switch-level congestion estimates produced by MLCong to dynamically determine the amount of inflation to apply to each switch to resolve congestion. The third model, DLRoute, is a convolutional neural network that uses the previous heat maps to predict whether or not a placement solution is routable. Once a placement solution is determined to be routable, further optimization may be avoided, leading to improved runtimes. Experimental results obtained using 372 benchmarks provided by Xilinx Inc. show that when all three models are integrated into GPlace3.0, placement runtimes decrease by an average of 48%.
Hannah Szentimrey, Abeer Alhyari, Jérémy Foxcroft, Timothy Martin, David Noel, Gary William Grewal, Shawki Areibi
ACM Trans. Design Autom. Electr. Syst.7
2019 A Flat Timing-Driven Placement Flow for Modern FPGAs
abstract
In this paper, we propose a novel, flat analytic timing-driven placer without explicit packing for Xilinx UltraScale FPGA devices. Our work uses novel methods to simultaneously optimize for timing, wirelength and congestion throughout the global and detailed placement stages. We evaluate the effectiveness of the flat placer on the ISPD 2016 benchmark suite for the xcvu095 UltraScale device, as well as on industrial benchmarks. Experimental results show that on average, FTPlace achieves an 8% increase in maximum clock rate, an 18% decrease in routed wirelength, and produces placements that require 80% less time to route when compared to Xilinx Vivado 2018.1.
Timothy Martin, Dani Maarouf, Ziad Abuowaimer, Abeer Alhyari, Gary William Grewal, Shawki Areibi
DAC6
2019 A Deep Learning Framework to Predict Routability for FPGA Circuit Placement
abstract
The ability to accurately and efficiently estimate the routability of a circuit based on its placement is one of the most challenging and difficult tasks in the Field Programmable Gate Array (FPGA) flow. In this paper, we present a novel, deep-learning framework based on a Convolutional Neural Network model for predicting the routability of a placement. We also incorporate the deep-learning model into a state-of-the-art placement tool, and show how the model can be used to (1) avoid costly, but futile, place-and-route iterations, and (2) improve the placer's ability to produce routable placements for hard-to-route circuits using feedback based on routability estimates generated by the proposed model. The model is trained and evaluated using over 26K placement images derived from 372 benchmarks supplied by Xilinx Inc. Experimental results show that the proposed framework achieves a routability prediction accuracy of 97%, while exhibiting runtimes of only a few milliseconds.
Abeer Alhyari, Ahmed Elshamli, Ziad Abuowaimer, Shawki Areibi, Gary William Grewal
FPL4
2019 Novel Congestion-estimation and Routability-prediction Methods based on Machine Learning for Modern FPGAs
abstract
Effectively estimating and managing congestion during placement can save substantial placement and routing runtime. In this article, we present a machine-learning model for accurately and efficiently estimating congestion during FPGA placement. Compared with the state-of-the-art machine-learning congestion-estimation model, our results show a 25% improvement in prediction accuracy. This makes our model competitive with congestion estimates produced using a global router. However, our model runs, on average, 291× faster than the global router. Overall, we are able to reduce placement runtimes by 17% and router runtimes by 19%. An additional machine-learning model is also presented that uses the output of the first congestion-estimation model to determine whether or not a placement is routable. This second model has an accuracy in the range of 93% to 98%, depending on the classification algorithm used to implement the learning model, and runtimes of a few milliseconds, thus making it suitable for inclusion in any placer with no worry of additional computational overhead.
Abeer Alhyari, Ziad Abuowaimer, Timothy Martin, Gary William Grewal, Shawki Areibi, Anthony Vannelli
ACM Trans. Reconfigurable Technol. Syst.5
2018 Machine-Learning Based Congestion Estimation for Modern FPGAs
abstract
Avoiding congestion for routing resources has become one of the most important placement objectives. In this paper, we present a machine-learning model for accurately and efficiently estimating congestion during FPGA placement. Compared with the state-of-the-art machine-learning congestion-estimation model, our results show a 25% improvement in prediction accuracy. This makes our model competitive with congestion estimates produced using a global router. However, our model runs, on average, 291x faster than the global router.
Dani Maarouf, Abeer Alhyari, Ziad Abuowaimer, Timothy Martin, Andrew David Gunter, Gary William Grewal, Shawki Areibi, Anthony Vannelli
FPL7
2018 Stochastic Layer-Wise Precision in Deep Neural Networks
Griffin Lacey, Graham W. Taylor, Shawki Areibi
UAI3
2018 GPlace3.0: Routability-Driven Analytic Placer for UltraScale FPGA Architectures
abstract
Optimizing for routability during FPGA placement is becoming increasingly important, as failure to spread and resolve congestion hotspots throughout the chip, especially in the case of large designs, may result in placements that either cannot be routed or that require the router to work excessively hard to obtain success. In this article, we introduce a new, analytic routability-aware placement algorithm for Xilinx UltraScale FPGA architectures. The proposed algorithm, called GPlace3.0, seeks to optimize both wirelength and routability. Our work contains several unique features including a novel window-based procedure for satisfying legality constraints in lieu of packing, an accurate congestion estimation method based on modifications to the pathfinder global router, and a novel detailed placement algorithm that optimizes both wirelength and external pin count. Experimental results show that compared to the top three winners at the recent ISPD’16 FPGA placement contest, GPlace3.0 is able to achieve (on average) a 7.53%, 15.15%, and 33.50% reduction in routed wirelength, respectively, while requiring less overall runtime. As well, an additional 360 benchmarks were provided directly from Xilinx Inc. These benchmarks were used to compare GPlace3.0 to the most recently improved versions of the first- and second-place contest winners. Subsequent experimental results show that GPlace3.0 is able to outperform the improved placers in a variety of areas including number of best solutions found, fewest number of benchmarks that cannot be routed, runtime required to perform placement, and runtime required to perform routing.
Ziad Abuowaimer, Dani Maarouf, Timothy Martin, Jérémy Foxcroft, Gary William Grewal, Shawki Areibi, Anthony Vannelli
ACM Trans. Design Autom. Electr. Syst.6
2017 A Machine Learning Framework for FPGA Placement (Abstract Only)
Gary William Grewal, Shawki Areibi, Matthew Westrik, Ziad Abuowaimer, Betty Zhao
FPGA2
2016 Caffeinated FPGAs: FPGA framework For Convolutional Neural Networks
abstract
Convolutional Neural Networks (CNNs) have gained significant traction in the field of machine learning, particularly due to their high accuracy in visual recognition. Recent works have pushed the performance of GPU implementations of CNNs showing significant improvements in their classification and training times. With these improvements, many frameworks have become available for implementing CNNs on both CPUs and GPUs, with no support for FPGA implementations. In this work we present a modified version of the popular CNN framework Caffe, with FPGA support. This allows for classification using CNN models and specialized FPGA implementations with the flexibility of reprogramming the device when necessary, seamless memory transactions between host and device, simple-to-use test benches, and the ability to create pipelined layer implementations. To validate the framework, we use the Xilinx SDAccel environment to implement an FPGA-based Winograd convolution engine and show that it can be used alongside other layers running on a host processor to run several popular CNNs (AlexNet, GoogleNet, VGG A, Overfeat). The results show that our framework achieves 50 GFLOPS across 3×3 convolutions in the benchmarks. This is achieved within a practical framework, which will aid in future development of FPGA-based CNNs.
Roberto DiCecco, Griffin Lacey, Jasmina Vasiljevic, Paul Chow, Graham W. Taylor, Shawki Areibi
FPT6
2016 GPlace: a congestion-aware placement tool for ultrascale FPGAs
abstract
Traditional FPGA flows that wait until the routing stage to tackle congestion are quickly becoming less effective. This is due to the increasing size and complexity of FPGA architectures and the designs targeted for them. In this paper, we present two new congestion-aware placement tools for Xilinx UltraScale architectures, called GPlace-pack and GPlace-flat, respectively. The former placer participated in the ISPD 2016 Routability-driven Placement Contest for FPGAs, and finished in third place overall. The latter placer was subseqently developed based on our experience in the contest with GPlace-pack. Results obtained indicate that GPlace-flat is on average 5.3× faster than GPlace-pack. The post routing results show that GPlace-flat is able to obtain a further 22.5% improvement in wirelength and a 40.0% improvement in runtime compared to GPlace-pack.
Ryan Pattison, Ziad Abuowaimer, Shawki Areibi, Gary William Grewal, Anthony Vannelli
ICCAD3
2012 A Dynamic Sampling Framework for Multi-class Imbalanced Data
abstract
In this paper we present a Dynamic Sampling Framework for use with multi-class imbalanced data containing any number of classes. The framework makes use of existing sampling techniques such as RUS, ROS, and SMOTE and ties the classification algorithm into the sampling process in a wrapper like manner. In doing so the framework is able to search for a desirably sampled training set, thus eliminating the need to specify a target distribution and automatically tuning the training set distribution to the classification algorithm's learning preferences. This is important when re-sampling multi-class data where manually searching for an appropriate target distribution would be a daunting task. We test both our Dynamic Sampling approach and traditional Static Sampling using RUS, ROS, SMOTE, ROS+RUS, and SMOTE+RUS with several classification algorithms on a four class, highly imbalanced data set. We compare the results of Static Sampling and Dynamic Sampling and find that overall both techniques are able to raise Recall for the highest minority classes, but Dynamic Sampling is also able to maintain or raise Recall for the majority classes. Also, Dynamic Sampling is overall more robust and resilient, and is better able to sustain classifier Accuracy and to raise G-Mean and Minimum F-Measures.
Bazyli Debowski, Shawki Areibi, Gary William Grewal, J. Tempelman
ICMLA (2)2
2012 A Sequential Ensemble Classification (SEC) System for Tackling the Problem of Unbalance Learning: A Case Study
abstract
In this paper we propose a Sequential Ensemble Classification (SEC) technique which is designed to tackle the problem of learning from a data set with an extremely unbalanced distribution of instances among the classes. This system employs a specific decomposition technique that reduces the degree of unbalance in the data by transforming multi-class problem into a sequence of binary class problems. We investigate two different implementations of the proposed method, one based on an ensemble of homogeneous classifiers and a second based on a heterogeneous ensemble of classifiers. A real-world medical data set has been chosen as a case study for the investigation of the proposed method. The data is highly unbalanced, consists of a wide range of class values, some of which contain only a few instances, and which is voluminous. Our experimental results show that both schemes of the SEC system are able to outperform standalone classifiers, with the highest performance being achieved by the homogeneous design of the system.
Samaneh Sheikh-Nia, Gary William Grewal, Shawki Areibi
ICMLA (2)3
2011 GBSA: A group based search algorithm for packet classification
abstract
Packet: classification is an ubiquitous and key building block for many critical network devices such as routing, fire-walls and load balancing. Despite the enormous number of research performed on this topic, yet it is still one of the main bottlenecks in designing fast network devices. In this paper we propose a novel algorithm GBSA for packet classification that is scalable, fast and efficient. On average the algorithm consumes 0.4 MB of memory for a 10k rule set. The classification time per packet in worst case is 2 μs, and the pre-processing speed is 3M Rule/sec based on a CPU operating at 3.4 GHz.
Omar Ahmed, Shawki Areibi
IWCMC2
2011 StarPlace: A new analytic method for FPGA placement
Gary William Grewal, Shawki Areibi
Integr.3
2010 Strength Pareto Particle Swarm Optimization and Hybrid EA-PSO for Multi-Objective Optimization
abstract
This paper proposes an efficient particle swarm optimization (PSO) technique that can handle multi-objective optimization problems. It is based on the strength Pareto approach originally used in evolutionary algorithms (EA). The proposed modified particle swarm algorithm is used to build three hybrid EA-PSO algorithms to solve different multi-objective optimization problems. This algorithm and its hybrid forms are tested using seven benchmarks from the literature and the results are compared to the strength Pareto evolutionary algorithm (SPEA2) and a competitive multi-objective PSO using several metrics. The proposed algorithm shows a slower convergence, compared to the other algorithms, but requires less CPU time. Combining PSO and evolutionary algorithms leads to superior hybrid algorithms that outperform SPEA2, the competitive multi-objective PSO (MO-PSO), and the proposed strength Pareto PSO based on different metrics.
Ahmed Elhossini, Shawki Areibi, Robert D. Dony
Evol. Comput.2
2010 Implementation Approaches Trade-Offs for WiMax OFDM Functions on Reconfigurable Platforms
abstract
This work investigates several approaches for implementing the OFDM functions of the fixed-WiMax standard on reconfigurable platforms. In the first phase, a custom RTL approach, using VHDL, is investigated. The approach shows the capability of a medium-size FPGA to accommodate the OFDM functions of a fixed-WiMax transceiver with only 50% occupation rate. In the second phase, a high-level approach based on the AccelDSP tool is used and compared to the custom RTL approach. The approach presents an easy flow to transfer MATLAB floating-point code into synthesizable cores. The AccelDSP approach shows an area overhead of 10%, while allowing early architectural exploration and accelerating the design time by a factor of two. However, the performance figure obtained is almost 1/4 of that obtained in the custom RTL approach. In the third phase, the Tensilica Xtensa configurable processor is targeted, which presents remarkable figures in terms of power, area, and design time. Comparing the three approaches indicates that the custom RTL approach has the lead in terms of performance. However, both the AccelDSP and the Tensilica Xtensa approaches show fast design time and early architectural exploration capability. In terms of power, the obtained estimation results show that the configurable Xtensa processor approach has the lead, where approximately the total power consumed is about 12--15 times less than those results obtained by the other two approaches.
Ahmad Sghaier, Shawki Areibi, Robert D. Dony
ACM Trans. Reconfigurable Technol. Syst.2
2008 IEEE802.16-2004 OFDM functions implementation on FPGAS with design exploration
abstract
The IEEE 802.16 standard, WiMAX, is a promising solution for broadband wireless access with implementations already happening. Those implementations are considering FPGAs as a viable option, because of their flexibility, performance, time-to-market and available resources. In this paper, two approaches are employed to map the IEEE 802.16-2004 standard on FPGAs: a pure VHDL and an AccelDSP-based approaches. Mapping the OFDM part of the digital baseband processor resulted an occupation rate of 25% and 30% resepctively. Timing results showed that the VHDL approach can obtain a better maximum frequency of operation of 171.8 MHz. However, the AccelDSP approach shows an advantage of a shorter design time and early trade-off analysis.
Ahmad Sghaier, Shawki Areibi, Robert D. Dony
FPL2
2007 Window Based Prototype Filter Design for Highly Oversampled Filter Banks in Audio Applications
abstract
This paper describes a window based method for designing near perfect-reconstruction prototype filters for highly oversampled, complex modulated filter banks. The design method extends some well-known simple methods for critically sampled filter banks to the over-sampled case and to the case of different length analysis and synthesis filters. This design method is simple and effective for designing a large range of filter bank configurations. The design method is particularly useful in developing audio applications using oversampled filter banks where the target system's requirements are highly variable. The simplicity and flexibility of the design method means that this one method can be used to generate multiple prototype filters as the application requirements change.
David Hermann, Edward Chau, Robert D. Dony, Shawki Areibi
ICASSP (2)4
2007 The Impact of Arithmetic Representation on Implementing MLP-BP on FPGAs: A Study
abstract
In this paper, arithmetic representations for implementing multilayer perceptrons trained using the error backpropagation algorithm (MLP-BP) neural networks on field-programmable gate arrays (FPGAs) are examined in detail. Both floating-point (FLP) and fixed-point (FXP) formats are studied and the effect of precision of representation and FPGA area requirements are considered. A generic very high-speed integrated circuit hardware description language (VHDL) program was developed to help experiment with a large number of formats and designs. The results show that an MLP-BP network uses less clock cycles and consumes less real estate when compiled in an FXP format, compared with a larger and slower functioning compilation in an FLP format with similar data representation width, in bits, or a similar precision and range.
Antony W. Savich, Medhat A. Moussa, Shawki Areibi
IEEE Trans. Neural Networks3
2004 An Island-Based GA Implementation for VLSI Standard-Cell Placement
Guangfa Lu, Shawki Areibi
GECCO (2)2
2004 Effective Memetic Algorithms for VLSI Design = Genetic Algorithms + Local Search + Multi-Level Clustering
abstract
Combining global and local search is a strategy used by many successful hybrid optimization approaches. Memetic Algorithms (MAs) are Evolutionary Algorithms (EAs) that apply some sort of local search to further improve the fitness of individuals in the population. Memetic Algorithms have been shown to be very effective in solving many hard combinatorial optimization problems. This paper provides a forum for identifying and exploring the key issues that affect the design and application of Memetic Algorithms. The approach combines a hierarchical design technique, Genetic Algorithms, constructive techniques and advanced local search to solve VLSI circuit layout in the form of circuit partitioning and placement. Results obtained indicate that Memetic Algorithms based on local search, clustering and good initial solutions improve solution quality on average by 35% for the VLSI circuit partitioning problem and 54% for the VLSI standard cell placement problem.
Shawki Areibi, Zhen Yang 0006
Evol. Comput.1
2003 Design and optimization of multithreshold CMOS (MTCMOS) circuits
abstract
Reducing power dissipation is one of the most important issues in very large scale integration design today. Scaling causes subthreshold leakage currents to become a large component of total power dissipation. Multithreshold technology has emerged as a promising technique to reduce leakage power. This paper presents several heuristic techniques for efficient gate clustering in multithreshold CMOS circuits by modeling the problem via bin-packing (BP) and set-partitioning (SP) techniques. The SP technique takes the circuit's routing complexity into consideration which is critical for deep submicron (DSM) implementations. By applying the techniques to six benchmarks to verify functionality, results obtained indicate that our proposed techniques can achieve on average 84% savings for leakage power and 12% savings for dynamic power. Furthermore, four hybrid clustering techniques that combine the BP and SP techniques to produce a more efficient solution are also devised. Ground bounce was also taken as a design parameter in the optimization problem. While accounting for noise, the proposed hybrid solution achieves on average 9% savings for dynamic power and 72% savings for leakage power dissipation at sufficient speeds and adequate noise margins.
Mohab Anis, Shawki Areibi, Mohamed I. Elmasry
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2002 Advanced P2P Architecture Using Autonomous Agents
Hooman Homayounfar, Fangju Wang, Shawki Areibi
CAINE3
2002 Hardware Implementation of Genetic Algorithms for VLSI Design
G. Koonar, Shawki Areibi, Medhat A. Moussa
CAINE2
2002 Feasibility of Floating-Point Arithmetic in FPGA based ANNs
Kristian R. Nichols, Medhat A. Moussa, Shawki Areibi
CAINE3
2002 Global Placement Techniques for VLSI Physical Design Automation
Zhen Yang 0006, Shawki Areibi
CAINE2
2002 Dynamic and leakage power reduction in MTCMOS circuits using an automated efficient gate clustering technique
abstract
Reducing power dissipation is one of the most principle subjects in VLSI design today. Scaling causes subthreshold leakage currents to become a large component of total power dissipation. This paper presents two techniques for efficient gate clustering in MTCMOS circuits by modeling the problem via Bin-Packing (BP) and Set-Partitioning (SP) techniques. An automated solution is presented, and both techniques are applied to six benchmarks to verify functionality. Both methodologies offer significant reduction in both dynamic and leakage power over previous techniques during the active and standby modes respectively. Furthermore, the SP technique takes the circuit's routing complexity into consideration which is critical for Deep Sub-Micron (DSM) implementations. Sufficient performance is achieved, while significantly reducing the overall sleep transistors' area. Results obtained indicate that our proposed techniques can achieve on average 90% savings for leakage power and 15% savings for dynamic power.
Mohab Anis, Mohamed I. Elmasry, Shawki Areibi
DAC4
2002 An Adaptive Genetic Algorithm For Multi Objective Flexible Manufacturing Systems
Abdulnasser Younes, Hamada H. Ghenniwa, Shawki Areibi
GECCO3
1999 Attractor-repeller approach for global placement
abstract
Traditionally, analytic placement has used linear or quadratic wirelength objective functions. Minimizing either formulation attracts cells sharing common signals (nets) together. The result is a placement with a great deal of overlap among the cells. To reduce cell overlap, the methodology iterates between global optimization and repartitioning of the placement area. In this work, we added new attractive and repulsive forces to the traditional formulation so that overlap among cells is diminished without repartitioning the placement area. The superiority of our approach stems from the fact that our new formulations are convex and no hard constraints are required. A preliminary version of the new placement method is tested using a set of MCNC benchmarks and, on average, the new method achieved 3.96% and 7.6% reduction in wirelength and CPU time compared to TimberWolf v7.0 in the hierarchical mode.
Hussein Etawil, Shawki Areibi, Anthony Vannelli
ICCAD2
1993 Circuit Partitioning Using a Tabu Search Approach
Shawki Areibi, Anthony Vannelli
ISCAS1