Tarek A. El-Ghazawi

dblp:e/TarekAElGhazawi · DBLP profile ↗
← Back
111ranked-venue papers
11as first author
7since 2021 · last 2025
0000-0001-9687-7939ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 85 · 11 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2Graphics, computer vision, multimedia, augmented reality and games · 2Software engineering, systems software and programming languages · 1Theory of computation · 1
YearPublicationVenuePosition
2025 QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration
abstract
The deployment of mixture-of-experts (MoE) large language models (LLMs) presents significant challenges due to their high memory demands. These challenges become even more pronounced in multi-tenant environments, where shared resources must accommodate multiple models, limiting the effectiveness of conventional virtualization techniques. This paper addresses the problem of efficiently serving multiple fine-tuned MoE-LLMs on a single GPU. We propose a serving system that employs similarity-based expert consolidation to reduce the overall memory footprint by sharing similar experts across models. To ensure output quality, we introduce runtime partial reconfiguration, dynamically replacing non-expert layers when processing requests from different models. As a result, our approach achieves competitive output quality while maintaining throughput comparable to serving a single model, and incurs only a negligible increase in time-to-first-token (TTFT). Experiments on a server with a single NVIDIA A100 GPU (80GB) using Mixtral-8x7B models demonstrate an 85% average reduction in turnaround time compared to NVIDIA’s multi-instance GPU (MIG). Furthermore, experiments on Google’s Switch Transformer Base-8 model with up to four variants demonstrate the scalability and resilience of our approach in maintaining output quality compared to other model merging baselines, highlighting its effectiveness.
Hamid Reza Imani, Peiman Mohseni, Abdolah Amirany, Tarek A. El-Ghazawi
ICML5
2023 Virtualizing a Post-Moore's Law Analog Mesh Processor: The Case of a Photonic PDE Accelerator
abstract
Innovative processor architectures aim to play a critical role in future sustainment of performance improvements under severe limitations imposed by the end of Moore’s Law. The Reconfigurable Optical Computer (ROC) is one such innovative, Post-Moore’s Law processor. ROC is designed to solve partial differential equations in one shot as opposed to existing solutions, which are based on costly iterative computations. This is achieved by leveraging physical properties of a mesh of optical components that behave analogously to lumped electrical components. However, virtualization is required to combat shortfalls of the accelerator hardware. Namely, (1) the infeasibility of building large photonic arrays to accommodate arbitrarily large problems and (2) underutilization brought about by mismatches in problem and accelerator mesh sizes due to future advances in manufacturing technology. In this work, we introduce an architecture and methodology for lightweight virtualization of ROC that exploits advantages borne from optical computing technology. Specifically, we apply temporal and spatial virtualization to ROC and then extend the accelerator scheduling tradespace with the introduction of spectral virtualization. Additionally, we investigate multiple resource scheduling strategies for a system-on-chip (SoC)-based PDE acceleration architecture and show that virtual configuration management offers a speedup of approximately 2×. Finally, we show that overhead from virtualization is minimal, and our experimental results show two orders of magnitude increased speed as compared to microprocessor execution while keeping errors due to virtualization under 10%.
Jeff Anderson, Engin Kayraklioglu, Hamid Reza Imani, Chen Shen 0005, Mario Miscuglio, Volker J. Sorger, Tarek A. El-Ghazawi
ACM Trans. Embed. Comput. Syst.7
2023 Online Service Migration in Mobile Edge With Incomplete System Information: A Deep Recurrent Actor-Critic Learning Approach
abstract
Multi-access Edge Computing (MEC) is an emerging computing paradigm that extends cloud computing to the network edge to support resource-intensive applications on mobile devices. As a crucial problem in MEC, service migration needs to decide how to migrate user services for maintaining the Quality-of-Service when users roam between MEC servers with limited coverage and capacity. However, finding an optimal migration policy is intractable due to the dynamic MEC environment and user mobility. Many existing studies make centralized migration decisions based on complete system-level information, which is time-consuming and also lacks desirable scalability. To address these challenges, we propose a novel learning-driven method, which is user-centric and can make effective online migration decisions by utilizing incomplete system-level information. Specifically, the service migration problem is modeled as a Partially Observable Markov Decision Process (POMDP). To solve the POMDP, we design a new encoder network that combines a Long Short-Term Memory (LSTM) and an embedding matrix for effective extraction of hidden information, and further propose a tailored off-policy actor-critic algorithm for efficient training. The extensive experimental results based on real-world mobility traces demonstrate that this new method consistently outperforms both the heuristic and state-of-the-art learning-driven algorithms and can achieve near-optimal results on various MEC scenarios.
Jin Wang 0024, Jia Hu 0001, Geyong Min, Qiang Ni, Tarek A. El-Ghazawi
IEEE Trans. Mob. Comput.5
2022 iSample: Intelligent Client Sampling in Federated Learning
abstract
The pervasiveness of AI in society has made machine learning (ML) an invaluable tool for mobile and internet-of-things (IoT) devices. While the aggregate amount of data yielded by those devices is sufficient for training an accurate model, the data available to any one device is limited. Therefore, augmenting the learning at any of the devices with the experience from observations associated with the rest of the devices will be necessary. This, however, can dramatically increase the bandwidth requirements. Prior work has led to the development of Federated Learning (FL), where instead of exchanging data, client devices can only share weights to learn from one another. However, het-erogeneity in device resource availability and network conditions still impose limitations on training performance. In order to improve performance while maintaining good levels of accuracy, we introduce iSample. iSample, an intelligent sampling technique, selects clients by jointly considering known network performance and model quality parameters, allowing the minimization of training time. We compare iSample with other federated learning approaches and show that iSample improves the performance of the global model, especially in the earlier stages of training, while decreasing the training time for both CNN and VGG by 27% and 39%, respectively.
Hamid Reza Imani, Jeff Anderson, Tarek A. El-Ghazawi
ICFEC3
2022 A Deep Neural Network Accelerator using Residue Arithmetic in a Hybrid Optoelectronic System
abstract
The acceleration of Deep Neural Networks (DNNs) has attracted much attention in research. Many critical real-time applications benefit from DNN accelerators but are limited by their compute-intensive nature. This work introduces an accelerator for Convolutional Neural Network (CNN) , based on a hybrid optoelectronic computing architecture and residue number system (RNS) . The RNS reduces the optical critical path and lowers the power requirements. In addition, the wavelength division multiplexing (WDM) allows high-speed operation at the system level by enabling high-level parallelism. The proposed RNS compute modules use one-hot encoding, and thus enable fast switching between the electrical and optical domains. We propose a new architecture that combines residue electrical adders and optical multipliers as the matrix-vector multiplication unit. Moreover, we enhance the implementation of different CNN computational kernels using WDM-enabled RNS based integrated photonics. The area and power efficiency of the proposed accelerator are 0.39 TOPS/s/mm 2 and 3.22 TOPS/s/W, respectively. In terms of computation capability, the proposed chip is 12.7× and 4.02× better than other optical implementation and memristor implementation, respectively. Our experimental evaluation using DNN benchmarks illustrates that our architecture can perform on average more than 72 times faster than GPU under the same power budget.
Yousra Al-Kabani, Krunal Puri, Volker J. Sorger, Tarek A. El-Ghazawi
ACM J. Emerg. Technol. Comput. Syst.6
2022 Adaptive and Efficient Resource Allocation in Cloud Datacenters Using Actor-Critic Deep Reinforcement Learning
abstract
The ever-expanding scale of cloud datacenters necessitates automated resource provisioning to best meet the requirements of low latency and high energy-efficiency. However, due to the dynamic system states and various user demands, efficient resource allocation in cloud faces huge challenges. Most of the existing solutions for cloud resource allocation cannot effectively handle the dynamic cloud environments because they depend on the prior knowledge of a cloud system, which may lead to excessive energy consumption and degraded Quality-of-Service (QoS). To address this problem, we propose an adaptive and efficient cloud resource allocation scheme based on Actor-Critic Deep Reinforcement Learning (DRL). First, the actor parameterizes the policy (allocating resources) and chooses actions (scheduling jobs) based on the scores assessed by the critic (evaluating actions). Next, the resource allocation policy is updated by using gradient ascent while the variance of policy gradient is reduced with an advantage function, which improves the training efficiency of the proposed method. We conduct extensive simulation experiments using real-world data from Google cloud datacenters. The results show that our method can obtain the superior QoS in terms of latency and job dismissing rate with enhanced energy-efficiency, compared to two advanced DRL-based and five classic cloud resource allocation methods.
Zheyi Chen, Jia Hu 0001, Geyong Min, Chunbo Luo, Tarek A. El-Ghazawi
IEEE Trans. Parallel Distributed Syst.5
2021 A Machine-Learning-Based Framework for Productive Locality Exploitation
abstract
Data locality is of extreme importance in programming distributed-memory architectures due to its implications on latency and energy consumption. Automated compiler and runtime system optimization studies have attempted to improve data locality exploitation without burdening the programmer. However, due to the difficulty of static code analysis, conservatism in compiler optimizations to avoid errors, and cost of dynamic analysis, the efficacy of automated optimizations is limited. Therefore, programmers need to spend significant effort in optimizing locality while creating applications for distributed memory parallel systems. We present a machine-learning based framework to automatically exploit locality in distributed memory applications. This framework takes application source whose time-critical blocks are marked by pragmas, and produces optimized source code that uses a regressor for efficient data movement. The regressor is trained with automatically-collected application profiles with very small input data sizes. We integrate our prototype in the Chapel language stack. In our experiments, we show that the Elastic Net model is the ideal regressor for our case and applications that utilize Elastic Net can perform very similarly to programmer-optimized versions. We also show that such regressors can be trained within few minutes on a cluster or within 30 minutes on a workstation, including data collection.
Engin Kayraklioglu, Erwan Favry, Tarek A. El-Ghazawi
IEEE Trans. Parallel Distributed Syst.3
2020 A Design Methodology for Post-Moore's Law Accelerators: The Case of a Photonic Neuromorphic Processor
abstract
Over the past decade alternative technologies have gained momentum as conventional digital electronics continue to approach their limitations, due to the end of Moore’s Law and Dennard Scaling. At the same time, we are facing new application challenges such as those due to the enormous increase in data. The attention, has therefore, shifted from homogeneous computing to specialized heterogeneous solutions. As an example, brain-inspired computing has re-emerged as a viable solution for many applications. Such new processors, however, have widened the abstraction gamut from device level to applications. Therefore, efficient abstractions that can provide vertical design-flow tools for such technologies became critical. Photonics in general, and neuromorphic photonics in particular, are among the promising alternatives to electronics. While the arsenal of device level toolbox for photonics, and high-level neural network platforms are rapidly expanding, there has not been much work to bridge this gap. Here, we present a design methodology to mitigate this problem by extending high-level hardware-agnostic neural network design tools with functional and performance models of photonic components. In this paper we detail this tool and methodology by using design examples and associated results. We show that adopting this approach enables designers to efficiently navigate the design space and devise hardware-aware systems with alternative technologies.
Armin Mehrabian, Volker J. Sorger, Tarek A. El-Ghazawi
ASAP3
2020 Software stack for an analog mesh computer: the case of a nanophotonic PDE accelerator
abstract
The slowing of Moore's Law is forcing the computer industry to embrace domain-specific hardware, which must be coupled with general-purpose traditional systems. This architecture is most useful when large compute power is needed. Among the most compute-intensive applications is the simulation of physical sciences. To maximize productivity in this domain, a variety accelerators have been proposed; however, the analog mesh computer has consistently been proven to require the shortest time-to-solution when targeted toward the Poisson equation. Recent advances in material science have increased the flexibility of the analog mesh computer, positioning it well for future heterogeneous computing systems. However, for the analog mesh computer to gain widespread acceptance, a software stack is required to enable seamless integration with a classical computer. Here, we introduce a software stack designed for the class of analog mesh computers that efficiently generates mesh mappings of a physical problem by enabling users to describe their problem in terms of boundary conditions and mesh parameters. Experiments on a specific implementation of analog mesh computer, the nanophotonic partial differential equation accelerator, show that this stack enables problem-to-mesh scalability expected by the scientific community.
Engin Kayraklioglu, Jeff Anderson, Hamid Reza Imani, Volker J. Sorger, Tarek A. El-Ghazawi
CF5
2020 DNNARA: A Deep Neural Network Accelerator using Residue Arithmetic and Integrated Photonics
abstract
Deep Neural Networks (DNNs) are currently used in many fields, including critical real-time applications. Due to its compute-intensive nature, speeding up DNNs has become an important topic in current research. We propose a hybrid opto-electronic computing architecture targeting the acceleration of DNNs based on the residue number system (RNS). In this novel architecture, we combine the use of Wavelength Division Multiplexing (WDM) and RNS for efficient execution. WDM is used to enable a high level of parallelism while reducing the number of optical components needed to decrease the area of the accelerator. Moreover, RNS is used to generate optical components with short optical critical paths. In addition to speed, this has the advantage of lowering the optical losses and reducing the need for high laser power. Our RNS compute modules use one-hot encoding and thus enable fast switching between the electrical and optical domains.
Yousra Al-Kabani, Volker J. Sorger, Tarek A. El-Ghazawi
ICPP5
2020 Communication and security in communicating things networks
Hicham Lakhlef, Julien Bourgeois, Saad Harous, Tarek A. El-Ghazawi
Ad Hoc Networks4
2020 Towards Accurate Prediction for High-Dimensional and Highly-Variable Cloud Workloads with Deep Learning
abstract
Resource provisioning for cloud computing necessitates the adaptive and accurate prediction of cloud workloads. However, the existing methods cannot effectively predict the high-dimensional and highly-variable cloud workloads. This results in resource wasting and inability to satisfy service level agreements (SLAs). Since recurrent neural network (RNN) is naturally suitable for sequential data analysis, it has been recently used to tackle the problem of workload prediction. However, RNN often performs poorly on learning long-term memory dependencies, and thus cannot make the accurate prediction of workloads. To address these important challenges, we propose a deep Learning based Prediction Algorithm for cloud Workloads (L-PAW). First, a top-sparse auto-encoder (TSA) is designed to effectively extract the essential representations of workloads from the original high-dimensional workload data. Next, we integrate TSA and gated recurrent unit (GRU) block into RNN to achieve the adaptive and accurate prediction for highly-variable workloads. Using real-world workload traces from Google and Alibaba cloud data centers and the DUX-based cluster, extensive experiments are conducted to demonstrate the effectiveness and adaptability of the L-PAW for different types of workloads with various prediction lengths. Moreover, the performance results show that the L-PAW achieves superior prediction accuracy compared to the classic RNN-based and other workload prediction methods for high-dimensional and highly-variable real-world cloud workloads.
Zheyi Chen, Jia Hu 0001, Geyong Min, Albert Y. Zomaya, Tarek A. El-Ghazawi
IEEE Trans. Parallel Distributed Syst.5
2019 Photonic Processor for Fully Discretized Neural Networks
abstract
Machine learning is now moving towards, and will become prevalent in, fog-computing and real-time computing environments. To this end, much machine-learning-at-the-edge research has focused on efficient neural network architectures, giving rise to efficient approximations of fixed-point neural networks, called discretized neural networks. While higher performing than their fixed and floating-point counterparts, discretized neural networks still have an existing bottleneck at the neuron's accumulation of products, called the popcount. This bottleneck sets an upper bound on performance regardless of neural network architecture. We address the popcount bottleneck by introducing a photonic discretized neural network processor. This processor minimizes the popcount bottleneck, thereby maximizing neural network computational throughput. Additionally, it offers potential for performance enhancement through simultaneous convolution operations enabled by wavelength division multiplexing. We show that the photonic architecture is capable of increasing performance by 700% and 100% when compared to state-of-the-art digital and analog architectures, respectively.
Jeff Anderson, Yousra Al-Kabani, Volker J. Sorger, Tarek A. El-Ghazawi
ASAP5
2019 A Machine Learning Approach for Productive Data Locality Exploitation in Parallel Computing Systems
abstract
Data locality is of extreme importance in programming distributed-memory architectures due to its implications on latency and energy consumption. Automated compiler and runtime system optimization studies have attempted to improve data locality exploitation without burdening the programmer. However, due to the difficulty of static code analysis, conservatism in compiler optimizations to avoid errors, and cost of dynamic analysis, the efficacy of automated optimizations is limited. Therefore, programmers need to spend significant effort in optimizing locality. In this work, we present an automated code optimization framework that trains neural networks using application profiles for small data sizes that exhibit similar patterns to larger cases. The application is then modified to use the neural network to improve data locality exploitation. We prototype our framework for the Chapel language and integrate with the language stack. We experimentally demonstrate that our framework can learn access patterns and create optimized executables in minutes. The resulting executables perform more than one order of magnitude faster than unoptimized code, and are comparable to manual locality optimization without burdening the programmer and hindering productivity.
Engin Kayraklioglu, Erwan Favry, Tarek A. El-Ghazawi
CCGRID3
2018 APAT: an access pattern analysis tool for distributed arrays
abstract
Distributed arrays reduce programming effort through implicit communication. However, relying solely on this abstraction causes fine-grained communication and performance overhead. A variety of optimization techniques can be used to mitigate such overheads. However, these techniques require a thorough understanding of how distributed arrays are accessed which can be very challenging in realistic use cases. We present Access Pattern Analysis Tool (APAT) for distributed arrays. APAT is a framework that can be integrated into language software stack to efficiently collect access logs and analyze them. We show that APAT can help discover optimization opportunities that can lead to up to 35% improvement.
Engin Kayraklioglu, Tarek A. El-Ghazawi
CF2
2018 D3NoC: a dynamic data-driven hybrid photonic plasmonic NoC
abstract
It was previously shown that Hybrid Photonic Plasmonic Interconnect (HyPPI) is an efficient candidate for augmenting electronic network on chips (NoCs). Here we introduce a reconfigurable Hybrid Photonic Plasmonic NoC termed D3NOC, which intelligently augments electrical meshes with a hybrid photon-plasmon interconnect express bus. The intelligence uses the Dynamic Data Driven Application System (DDDAS) paradigm, where computations and measurements form a dynamic closed feedback loop. Our results show up to 67% latency improvements and 69% dynamic power net improvements beyond overhead-corrected performance compared to a 16 × 16 base electrical mesh.
Armin Mehrabian, Vikram K. Narayana, Jeff Anderson, Volker J. Sorger, Tarek A. El-Ghazawi
CF7
2018 LAPPS: Locality-Aware Productive Prefetching Support for PGAS
abstract
Prefetching is a well-known technique to mitigate scalability challenges in the Partitioned Global Address Space (PGAS) model. It has been studied as either an automated compiler optimization or a manual programmer optimization. Using the PGAS locality awareness, we define a hybrid tradeoff. Specifically, we introduce locality-aware productive prefetching support for PGAS. Our novel, user-driven approach strikes a balance between the ease-of-use of compiler-based automated prefetching and the high performance of the laborious manual prefetching. Our prototype implementation in Chapel shows that significant scalability and performance improvements can be achieved with minimal effort in common applications.
Engin Kayraklioglu, Michael P. Ferguson, Tarek A. El-Ghazawi
ACM Trans. Archit. Code Optim.3
2017 HPC-Oriented Toolchain for Hardware Simulators
abstract
Hardware design is an essential part of research in high performance computing. Initial efforts in hardware research consist of analyzing the design ideas in a software simulator. This allows chip designers to minimize amount of manufacturing that would be too costly and to avoid doing FPGA designs which are even more time consuming. Simulating a hardware design involves running many tests that try different configurations. Moreover, hardware simulators generally do not support multi-threaded simulation. This causes major scalability issues as simulated HPC architectures have increasing number of cores.In this paper, we present a front-end framework for hardware simulators that allows chip designers to create simulation recipes and run them in parallel. This way, a cluster can easily be used to parallelize the hardware simulations. Our framework is implemented in Python3 and have functions such as running unlimited configurations, cooperating with job managers such as Slurm and SGE and collecting and parsing results.
Olivier Serres, Engin Kayraklioglu, Tarek A. El-Ghazawi
CLUSTER3
2017 HyPPI NoC: Bringing Hybrid Plasmonics to an Opto-Electronic Network-on-Chip
abstract
As we move towards an era of hundreds of cores, the research community has witnessed the emergence of optoelectronic network on-chip designs based on nanophotonics, in order to achieve higher network throughput, lower latencies, and lower dynamic power. However, traditional nanophotonics options face limitations such as large device footprints compared with electronics, higher static power due to continuous laser operation, and an upper limit on achievable data rates due to large device capacitances. Nanoplasmonics is an emerging technology that has the potential for providing transformative gains on multiple metrics due to its potential to increase the light-matter interaction. In this paper, we propose and analyze a hybrid opto-electric NoC that incorporates Hybrid Plasmonics Photonics Interconnect (HyPPI), an optical interconnect that combines photonics with plasmonics. We explore various opto-electronic network hybridization options by augmenting a mesh network with HyPPI links, and compare them with the equivalent options afforded by conventional nanophotonics as well as pure electronics. Our design space exploration indicates that augmenting an electronic NoC with HyPPI gives a performance to cost ratio improvement of up to 1.8×. To further validate our estimates, we conduct trace based simulations using the NAS Parallel Benchmark suite. These benchmarks show latency improvements up to 1.64×, with negligible energy increase. We then further carry out performance and cost projections for fully optical NoCs, using HyPPI as well as conventional nanophotonics. These futuristic projections indicate that all-HyPPI NoCs would be two orders more energy efficient than electronics, and two orders more area efficient than all-photonic NoCs.
Vikram K. Narayana, Armin Mehrabian, Volker J. Sorger, Tarek A. El-Ghazawi
ICPP5
2017 Optimizing thin client caches for mobile cloud computing: : Design space exploration using genetic algorithms
abstract
Summary The emergence and rapid spread of interest and use of cloud computing as an accessible and expandable, as needed, computing facility on the go, has a very deep affinity to the proliferation of intelligent mobile devices including smartphones and tablets. Together, these technologies have the potential of not leaving anybody behind when it comes to computing applications whether small and personal or large and organizational, and regardless of geographic boundaries and economical conditions. However, many technical challenges still exist that are still delaying the realization of this dream with the responsiveness and quality needed from the user perspective. In this paper, we examine user requirements for access to the cloud through thin clients, handheld and mobile devices. In light of these requirements we characterize some of the needed research developments particularly in the area of device architecture. We present our work in exploring the cache design space for embedded processors using evolutionary techniques for mobile and thin client processors. We present a heuristic, evolutionary approach (genetic algorithm) to exploration that significantly cuts down on the time and resources, obtaining a near optimal design. We demonstrate the real‐world utility of our tool‐chain—“CERE” (pronounced SIRI) short for (CachE Recommendation Engine)—by rapidly and efficiently designing a cache hierarchy, which maximizes the performance of a web browser navigating to a set of popular websites running on a single ARM core. The goal is to improve the users' experience using web browsers. “CERE” made the right choices, and we were able to observe a 17.1%speedup going from the “best” hierarchy relative to the “worst” hierarchy. We will detail potential future directions as well.
Abdel-Hameed A. Badawy, Gabriel Yessin, Vikram K. Narayana, David Mayhew, Tarek A. El-Ghazawi
Concurr. Comput. Pract. Exp.5
2016 Novel Models of Visual Topographic Map Alignment in the Superior Colliculus
abstract
The establishment of precise neuronal connectivity during development is critical for sensing the external environment and informing appropriate behavioral responses. In the visual system, many connections are organized topographically, which preserves the spatial order of the visual scene. The superior colliculus (SC) is a midbrain nucleus that integrates visual inputs from the retina and primary visual cortex (V1) to regulate goal-directed eye movements. In the SC, topographically organized inputs from the retina and V1 must be aligned to facilitate integration. Previously, we showed that retinal input instructs the alignment of V1 inputs in the SC in a manner dependent on spontaneous neuronal activity; however, the mechanism of activity-dependent instruction remains unclear. To begin to address this gap, we developed two novel computational models of visual map alignment in the SC that incorporate distinct activity-dependent components. First, a Correlational Model assumes that V1 inputs achieve alignment with established retinal inputs through simple correlative firing mechanisms. A second Integrational Model assumes that V1 inputs contribute to the firing of SC neurons during alignment. Both models accurately replicate in vivo findings in wild type, transgenic and combination mutant mouse models, suggesting either activity-dependent mechanism is plausible. In silico experiments reveal distinct behaviors in response to weakening retinal drive, providing insight into the nature of the system governing map alignment depending on the activity-dependent strategy utilized. Overall, we describe novel computational frameworks of visual map alignment that accurately model many aspects of the in vivo process and propose experiments to test them.
Ruben A. Tikidji-Hamburyan, Tarek A. El-Ghazawi, Jason W. Triplett
PLoS Comput. Biol.2
2016 Exploiting Hierarchical Locality in Deep Parallel Architectures
abstract
Parallel computers are becoming deeply hierarchical. Locality-aware programming models allow programmers to control locality at one level through establishing affinity between data and executing activities. This, however, does not enable locality exploitation at other levels. Therefore, we must conceive an efficient abstraction of hierarchical locality and develop techniques to exploit it. Techniques applied directly by programmers, beyond the first level, burden the programmer and hinder productivity. In this article, we propose the Parallel Hierarchical Locality Abstraction Model for Execution (PHLAME). PHLAME is an execution model to abstract and exploit machine hierarchical properties through locality-aware programming and a runtime that takes into account machine characteristics, as well as a data sharing and communication profile of the underlying application. This article presents and experiments with concepts and techniques that can drive such runtime system in support of PHLAME. Our experiments show that our techniques scale up and achieve performance gains of up to 88%.
Ahmad Anbar, Olivier Serres, Engin Kayraklioglu, Abdel-Hameed A. Badawy, Tarek A. El-Ghazawi
ACM Trans. Archit. Code Optim.5
2016 Enabling PGAS Productivity with Hardware Support for Shared Address Mapping: A UPC Case Study
abstract
Due to its rich memory model, the partitioned global address space (PGAS) parallel programming model strikes a balance between locality-awareness and the ease of use of the global address space model. Although locality-awareness can lead to high performance, supporting the PGAS memory model is associated with penalties that can hinder PGAS’s potential for scalability and speed of execution. This is because mapping the PGAS memory model to the underlying system requires a mapping process that is done in software, thereby introducing substantial overhead for shared accesses even when they are local. Compiler optimizations have not been sufficient to offset this overhead. On the other hand, manual code optimizations can help, but this eliminates the productivity edge of PGAS. This article proposes a processor microarchitecture extension that can perform such address mapping in hardware with nearly no performance overhead. These extensions are then availed to compilers through extensions to the processor instructions. Thus, the need for manual optimizations is eliminated and the productivity of PGAS languages is unleashed. Using Unified Parallel C (UPC), a PGAS language, we present a case study of a prototype compiler and architecture support. Two different implementations of the system were realized. The first uses a full-system simulator, gem5, which evaluates the overall performance gain of the new hardware support. The second uses an FPGA Leon3 soft-core processor to verify implementation feasibility and to parameterize the cost of the new hardware. The new instructions show promising results on all tested codes, including the NAS Parallel Benchmark kernels in UPC. Performance improvements of up to 5.5× for unmodified codes, sometimes surpassing hand-optimized performance, were demonstrated. We also show that our four-core FPGA prototype requires less than 2.4% of the overall chip’s area.
Olivier Serres, Abdullah Kayi, Ahmad Anbar, Tarek A. El-Ghazawi
ACM Trans. Archit. Code Optim.4
2015 Assessing Memory Access Performance of Chapel through Synthetic Benchmarks
abstract
The Partitioned Global Address Space(PGAS) programming model strikes a balance between high performance and locality awareness. As a PGAS language, Chapel relieves programmers from handling details of data movement in a distributed memory environment, by presenting a flat memory space that is logically partitioned among executing entities. Traversing such a space requires address mapping to the system virtual address space, and as such, this abstraction inevitably causes major overheads during memory accesses. In this paper, we analyzed the extent of this overhead by implementing a micro benchmark to test different types of memory accesses that can be observed in Chapel. We showed that, as the locality gets exploited speedup gains up to 35x can be achieved. This was demonstrated through hand tuning, however. More productive means should be provided to deliver such performance improvement without excessively burdening programmers. Therefore, we also discuss possibilities to increase Chapel's performance through standard libraries, compiler, runtime and/or hardware support to handle different types of memory accesses more efficiently.
Engin Kayraklioglu, Tarek A. El-Ghazawi
CCGRID2
2015 A Power-Aware Symbiotic Scheduling Algorithm for Concurrent GPU Kernels
abstract
The past several years have witnessed significant performance improvements in High-Performance Computing (HPC), due to the incorporation of GPUs as co-processors. On one hand, GPU devices are growing significantly in terms of the available number of cores and the memory hierarchy; as a result, effective utilization of the available GPU resources while limiting the system power consumption has become an issue of rising importance. On the other hand, GPU vendors are providing additional supporting features to make this easier, such as enabling concurrent execution of multiple kernels, and providing on-board power sensors that can accessed through software. Amidst these new developments, we are faced with new opportunities for efficiently scheduling GPU computational kernels under performance and power constraints. In this paper, we propose a power-aware scheduling technique that carries out both performance and power optimizations for concurrent GPU kernels. We have observed that for GPU kernels that are deployed for concurrent execution, the order in which the programmer specifies their invocation can significantly alter the execution time and the power draw. We attribute this behavior to the relative synergy (or lack thereof) among kernels that are launched within close proximity of each other. Accordingly, we define performance metrics for computing the extent to which kernels are symbiotic, as well as power metrics for reducing the overall power consumption. Both metrics are estimated by modeling the kernels' complementary resource requirements and execution characteristics. We then propose a power-aware symbiotic scheduling algorithm to obtain a concurrent kernel launch schedule with improved performance and reduced power consumption. Experimental studies are conducted on the Cray XK7 supercomputer with an NVIDIA K20 GPU in each node. The results demonstrate the efficacy of the proposed algorithm-based approach, which can be readily adopted by programmers with minimal programming effort and risk.
Teng Li 0009, Vikram K. Narayana, Tarek A. El-Ghazawi
ICPADS3
2015 Optimization of selected remote sensing algorithms for embedded Nvidia Kepler GPU architecture
abstract
This paper evaluates the potential of embedded Graphic Processing Units in the Nvidia's Tegra K1 for onboard processing. The performance is compared to a general purpose multi-core CPU and full fledge GPU accelerator. This study uses two algorithms: Wavelet Spectral Dimension Reduction of Hyperspectral Imagery and Automated Cloud-Cover Assessment (ACCA) Algorithm. Tegra K1 achieved 51% for ACCA algorithm and 20% for the dimension reduction algorithm, as compared to the performance of the high-end 8-core server Intel Xeon CPU with 13.5 times higher power consumption.
Lubomir Riha, Jacqueline LeMoigne-Stewart, Tarek A. El-Ghazawi
IGARSS3
2015 Communication efficient work distributions in stencil operation based applications
abstract
Summary In recent years, the use of accelerators in conjunction with CPUs, known as heterogeneous computing, has brought about significant performance increases for scientific applications. One of the best examples of this is lattice quantum chromodynamics (QCD), a stencil operation based simulation. These simulations have a large memory footprint necessitating the use of many graphics processing units (GPUs) in parallel. This requires the use of a heterogeneous cluster with one or more GPUs per node. In order to obtain optimal performance, it is necessary to determine an efficient communication pattern between GPUs on the same node and between nodes. In this paper, we present a performance model based method for minimizing the communication time of applications with stencil operations, such as lattice QCD, on heterogeneous computing systems with a non‐blocking InfiniBand interconnection network. The proposed method is able to increase the performance of the most computationally intensive kernel of lattice QCD by 25% due to improved overlapping of communication and computation. We also demonstrate that the aforementioned performance model and efficient communication patterns can be used to determine a cost efficient heterogeneous system design for stencil operation based applications. Copyright © 2014 John Wiley & Sons, Ltd.
Joseph Schneible, Lubomir Riha, Maria Malik, Tarek A. El-Ghazawi, Andrei Alexandru
Concurr. Comput. Pract. Exp.4
2015 Adaptive Cache Coherence Mechanisms with Producer-Consumer Sharing Optimization for Chip Multiprocessors
abstract
In chip multiprocessors (CMPs), maintaining cache coherence can account for a major performance overhead. Write-invalidate protocols adapted by most CMPs generate high cache-to-cache misses under producer–consumer sharing patterns. Accordingly, this paper presents three cache coherence mechanisms optimized for CMPs. First, to reduce coherence misses observed in write-invalidate-based protocols, we propose a dynamic write-update mechanism augmented on top of a write-invalidate protocol. This mechanism is specifically triggered at the detection of a producer–consumer sharing pattern. Second, we extend this adaptive protocol with a bandwidth-adaptive mechanism to eliminate performance degradation from write-updates under limited bandwidth. Finally, proximity-aware mechanism is proposed to extend the base adaptive protocol with latency-based optimizations. Experimental analysis is conducted on a set of scientific applications from the SPLASH-2 and NAS parallel benchmark suites. The proposed mechanisms were shown to reduce coherence misses by up to 48% and in return speed up application performance up to 30%. Bandwidth-adaptive mechanism is proven to perform well under varying levels of available bandwidth. Results from our proposed proximity-aware extension demonstrated up to 6% performance gain over the base adaptive protocol for 64-core tiled CMP runs. In addition, the analytical model provided good estimates for performance gains from our adaptive protocols.
Abdullah Kayi, Olivier Serres, Tarek A. El-Ghazawi
IEEE Trans. Computers3
2014 Predicting the severity of motor neuron disease progression using electronic health record data with a cloud computing Big Data approach
abstract
Motor neuron diseases (MNDs) are a class of progressive neurological diseases that damage the motor neurons. An accurate diagnosis is important for the treatment of patients with MNDs because there is no standard cure for the MNDs. However, the rates of false positive and false negative diagnoses are still very high in this class of diseases. In the case of Amyotrophic Lateral Sclerosis (ALS), current estimates indicate 10% of diagnoses are false-positives, while 44% appear to be false negatives. In this study, we developed a new methodology to profile specific medical information from patient medical records for predicting the progression of motor neuron diseases. We implemented a system using Hbase and the Random forest classifier of Apache Mahout to profile medical records provided by the Pooled Resource Open-Access ALS Clinical Trials Database (PRO-ACT) site, and we achieved 66% accuracy in the prediction of ALS progress.
Kyung Dae Ko, Tarek A. El-Ghazawi, Dongkyu Kim, Hiroki Morizono
CIBCB2
2014 Where should the threads go? Leveraging hierarchical data locality to solve the thread affinity dilemma
abstract
We are proposing a novel framework that amelio-rates locality-aware parallel programming models, by defining a hierarchical data locality model extension. We also propose two hierarchical thread partitioning algorithms. These algorithms synthesize hierarchical thread placement layouts that targets minimizing the program's overall communication costs. We demonstrate the effectiveness of our approach using the NAS Parallel Benchmarks implemented in Unified Parallel C (UPC) using a modified Berkeley UPC Compiler and runtime system. We achieved performance gains of up to 88% in performance by applying the placement layouts our algorithms suggest.
Ahmad Anbar, Abdel-Hameed A. Badawy, Olivier Serres, Tarek A. El-Ghazawi
ICPADS4
2013 Predictive energy management techniques for PGAS programming
abstract
Power consumption increasingly presents an upper bound on sustainable large scale computing performance and reliability. The Partitioned Global Address Space (PGAS) programming model is a family of parallel programming paradigms with a global address space for ease-of-use while providing locality awareness for efficient execution. Very little exploration has been done to determine the potential of PGAS programming models in improving scalable energy efficient computation for high performance computing (HPC) clusters. This paper examines features of the PGAS programming model that may support predictively reducing power consumption in distributed clusters via dynamic voltage frequency scaling (DVFS). These concepts are tested with Unified Parallel C (UPC) codes running on a cluster of commodity PCs which have been instrumented to measure power at the CPU socket level. We have also explored approaches to automating these power optimization techniques at compile time. Benchmarking results show a tangible reduction in power consumption without impacting the overall execution time of the program.
David K. Newsom, Sardar F. Azari, Ahmad Anbar, Tarek A. El-Ghazawi
AICCSA4
2013 Application-specific processors for web-browsing: An exploration and evaluation of the design space
abstract
The current trend in computing has been to add more and more to the CPU; especially bigger and bigger caches and more cache levels. Based on these observations, we sought to see if bigger is always better. We test this by performing an architectural design space exploration of various cache and frequency configurations for ARM processors. Analyzing the data, we made the surprising discovery that bigger is not always better and we should in fact be taking a step back in the architectural evolutionary roadmap for some applications. In this study, we performed an analysis of the performance of web-browsers versus the architectural configuration and related it to end-user satisfaction. In the end, we were able to determine that a scaled back modern core would not only be sufficient, but improve the performance of the web-browser. In doing this, we have also developed GW-GEM5 a set of tools for the creation, monitoring and analysis of concurrent gem5 simulations on computer clusters for use in design space parameter studies.
Gabriel Yessin, Lubomir Riha, Tarek A. El-Ghazawi, David Mayhew
ASAP3
2013 System architecture of the Mediterranean Dialogue Earth Observatory
abstract
A Mediterranean Dialogue Earth Observatory project is under implementation in Morocco. It is led by a group of experts from Turkey, Morocco and the USA. The objective of the project is to facilitate early warning and mitigation of a wide range of biogenic and anthropogenic disasters using remote sensing techniques. Some examples are flooding, storms, forest fires, climate change, recent public health incidents, such as malaria and avian influenza. The observatory comprises a network of real-time satellite remote sensing ground stations, installed at two universities in Morocco, Abdelmalek Essadi and AlAkhawayn, with a geostationary system to collect data from Meteosat geostationary satellite and a tracking station for polar orbiting satellites. The infrastructure includes as well the post processing computer clusters and relevant storage, software and distributional network. Archival and real-time remote sensing utilizing high performance computing clusters, are planned throughout the life cycle of disaster management. In this paper, a description of ground stations will be given, as well as the antennas connectivity and the access to data from partnering universities and collaborating end-users.
Chaker El Amrani, Gilbert Rochon, Tarek A. El-Ghazawi, Gulay Altay, Tajje-eddine Rachidi
IGARSS3
2012 A Compartive Study of Cloud Computing Middleware
abstract
Cloud computing is an emerging IT technology that is being used increasingly in industry, government and academia. There are several Cloud Computing middleware solutions available in the market. This paper proposes an approach and set of characteristics and metrics for comparing Cloud computing middleware based on functionality. Three popular open source middleware: Nimbus, Eucalyptus and Open Nebula, are also analyzed and evaluated on the basis of the proposed parameters. Future work will include more systems and an experimental benchmarking to study the relative performance.
Chaker El Amrani, Kaoutar Bahri Filali, Kaoutar Ben Ahmed, Amadou Tidiane Diallo, Stéphano Telolahy, Tarek A. El-Ghazawi
CCGRID6
2012 Distributed Shared Memory Programming in the Cloud
abstract
Cloud computing is beginning to play a dominant role in scientific computing. However, there are still several challenges that need to be addressed, before data-intensive scientific applications make the transition to the cloud. Adoption of a distributed shared memory (DSM) programming paradigm will be one approach to ease the transition, through the use of Partitioned Global Address Space (PGAS) languages. This paper explores initial results from the adoption of a PGAS language, Unified Parallel C, in programming a representative private cloud based on Eucalyptus.
Ahmad Anbar, Vikram K. Narayana, Tarek A. El-Ghazawi
CCGRID3
2012 Development of a real-time urban remote sensing initiative in the mediterranean region for early warning and mitigation of disasters
abstract
This project brings together a group of world class experts from research partner institutions in three countries: Turkey, Morocco and the USA, to plan and implement the North Atlantic Treaty Organization (NATO) Science for Peace sponsored Mediterranean Dialogue Earth Observatory (MDEO). The observatory comprises a network of real-time satellite remote sensing ground stations, to be established in Morocco. This investigation will also include a networked geostationary receiving station for the European Space Agency's Meteosat. The primary objective of the project is to facilitate early warning and mitigation of a wide range of biogenic and anthropogenic disasters. The project will also address mitigation of epidemics and epizootics, through identification and monitoring of infectious disease vector and reservoir habitat. Some examples of common concern among participating countries are flooding, storms, forest fires, climate change and its impacts, land use problems in agriculture, recent public health incidents, such as malaria, avian influenza, swine flu, as well as oil and hazardous chemical spills along the seashores [1]. Archival and real-time remote sensing and generation of near-real-time spatial data products, utilizing high performance computing clusters [2, 3], are planned throughout the life cycle of disaster management, including vulnerability assessment, infrastructure safeguards, early warning, emergency response, humanitarian relief, as well as post-disaster damage assessment, reconstruction and societal recovery [4, 5].
Chaker El Amrani, Gilbert Rochon, Tarek A. El-Ghazawi, Gulay Altay, Tajje-eddine Rachidi
IGARSS3
2012 Bandwidth Adaptive Write-update Optimizations for Chip Multiprocessors
abstract
Chip Multiprocessors (CMPs) have different technological parameters and physical constraints than earlier multi-processor systems, which should be taken into consideration when designing cache coherence protocols. Also, contemporary cache coherence protocols use invalidate schemes that are known to generate a high number of coherence misses. This is especially true under producer-consumer sharing patterns that can become a performance bottleneck as the number of cores increases. This paper presents two mechanisms to design efficient and scalable cache coherence protocols for CMPs. First, we propose an adaptive hybrid protocol to reduce coherence misses observed in write-invalidate based protocols. The proposed protocol is based on a write-invalidate scheme. However, adaptively, it can push updates to potential consumers based on observed producer- consumer sharing patterns. Secondly, we extend this adaptive protocol with an interconnection resource aware mechanism. Experimental evaluations, conducted on a tiled-CMP via full- system simulation, were used to assess the performance from our proposed dynamic hybrid protocols. Performance analysis is presented on a set of scientific applications from the SPLASH- 2 and NAS parallel benchmark suites. Results showed that the proposed mechanisms reduce cache-to-cache sharing misses up to 48% and in return speed up application performance up to 25%. In addition, the proposed interconnection resource aware mechanism is proven to perform well under varying interconnection utilizations.
Abdullah Kayi, Olivier Serres, Tarek A. El-Ghazawi
ISPA3
2012 Productivity of GPUs under different programming paradigms
abstract
SUMMARY Graphical processing units have been gaining rising attention because of their high performance processing capabilities for many scientific and engineering applications. However, programming such highly parallel devices requires adequate programming tools. Many such programming tools have emerged and hold the promise for high levels of performance. Some of such tools may require specialized parallel programming skills, while others attempt to target the domain scientist. The costs versus the benefits out of such tools are often unclear. In this work we examine the use of several of these programming tools such as Compute Unified Device Architecture, Open Compute Language, Portland Group Inc., and MATLAB in developing kernels from the (NAS) NASA Advanced Supercomputing parallel benchmarking suite. The resulting performance as well as the needed programmers' efforts were quantified and used to characterize the productivity of graphical processing units using these different programming paradigms. Copyright © 2011 John Wiley & Sons, Ltd.
Maria Malik, Teng Li 0009, Umar Sharif, Rabia Shahid, Tarek A. El-Ghazawi, Gregory B. Newby
Concurr. Comput. Pract. Exp.5
2012 Efficient Mapping of Task Graphs onto Reconfigurable Hardware Using Architectural Variants
abstract
High-performance reconfigurable computing involves acceleration of significant portions of an application using reconfigurable hardware. Mapping application task graphs onto reconfigurable hardware is, therefore, of rising attention. In this work, we approach the mapping problem by incorporating multiple architectural variants for each hardware task; the variants reflect tradeoffs between the logic resources consumed and the task execution throughput. We propose a mapping approach based on the genetic algorithm, and show its effectiveness for random task graphs as well as an N-body simulation application, demonstrating improvements of up to 78.6 percent in the execution time compared with choosing a fixed implementation variant for all tasks. We then validate our methodology through experiments on real hardware, an SRC-6 reconfigurable computer.
Miaoqing Huang, Vikram K. Narayana, Mohamed Bakhouya, Jaafar Gaber, Tarek A. El-Ghazawi
IEEE Trans. Computers5
2011 Modelling the performance of an SSD-Aware storage system using least squares regression
abstract
Flash memory has lately been used as a cache located between the system main memory and the magnetic hard drives in order to create robust and cost effective hybrid storage systems. The reason comes from the growing density of Solid State Devices (SSDs) at lower prices with main advantage of high random read efficiency compared to magnetic hard drives. When predicting the performance of such hybrid storage systems, it is inevitable to study the trade-off in selecting the different storage elements such as the main memory and the SSD cache relative to the capacity of the magnetic hard drive. The parameters of such prediction model are determined based on the application that the storage system would serve. In this paper, a prediction model that uses experimental evaluation of a hybrid storage system and real applications/benchmarks is used. The model utilizes parameters of both the storage system and applications in order to predict system performance based on metrics that are commonly used in storage system evaluation. The model allows the designer to select the best hybrid storage system parameters that satisfy certain application performance requirements. The model is highly accurate with a minimal error and a high prediction confidence level (95%) in reference to the experimental data collected from real applications using the proposed SSD-aware hybrid storage system.
Abdullah Aldahlawi, Esam El-Araby, Suboh A. Suboh, Tarek A. El-Ghazawi
AICCSA4
2011 Reflex Barrier: A Scalable Network-Based Synchronization Barrier
abstract
High-performance computing is witnessing the proliferation of multi-core processors in parallel architectures, and the trend is expected to increase further with the emerging many-core technology, leading to hundreds of processing cores within each compute node in the near future. Along side with this trend, it is also clear that total number of cores within the whole system is increasing. To be able to harvest the fruits of this massive parallelism, inter-process synchronization and communication should be as lightweight as they can be, and should be relying on as limited involvement as possible of the participating processors/cores. The synchronization algorithms that target shared memory processors are expected not to be able to scale on many-cores as they rely on atomics, locks, and/or cache coherence protocols, which all should be very costly operations on many-cores. In the same time, some many core architectures provide user space networks on chip (NoCs) that operate similar to regular networks. In this paper, we are introducing the Reflex barrier, a new synchronization barrier algorithm that relies on fundamental networking concepts. As the barrier relies on the characteristics of the network, it requires very little intervention from the participating processors/cores. The algorithm can also be implemented as split phase, which furnish an opportunity to reduce the synchronization cost. We implemented the algorithm using Unified Parallel C (UPC), MPI and pThreads. We tested our implementation on TILE64, a 64-core processor. The performance of the Reflex barrier is also analyzed and compared to other algorithms using performance models.
Ahmad Anbar, Olivier Serres, Tarek A. El-Ghazawi
ICPADS3
2011 A Static Task Scheduling Framework for Independent Tasks Accelerated Using a Shared Graphics Processing Unit
abstract
The High Performance Computing (HPC) field is witnessing the increasing use of Graphics Processing Units (GPUs) as application accelerators, due to their massively data-parallel computing architectures and exceptional floating-point computational capabilities. The performance advantage from GPU-based acceleration is primarily derived for GPU computational kernels that operate on large amount of data, consuming all of the available GPU resources. For applications that consist of several independent computational tasks that do not occupy the entire GPU, sequentially using the GPU one task at a time leads to performance inefficiencies. It is therefore important for the programmer to cluster small tasks together for sharing the GPU, however, the best performance cannot be achieved through an ad-hoc grouping and execution of these tasks. In this paper, we explore the problem of GPU tasks scheduling, to allow multiple tasks to efficiently share and be executed in parallel on the GPU. We analyze factors affecting multi-tasking parallelism and performance, followed by developing the multi-tasking execution model as a performance prediction approach. The model is validated by comparing with actual execution scenarios for GPU sharing. We then present the scheduling technique and algorithm based on the proposed model, followed by experimental verifications of the proposed approach using an NVIDIA Fermi GPU computing node. Our results demonstrate significant performance improvements using the proposed scheduling approach, compared with sequential execution of the tasks under the conventional multi-tasking execution scenario.
Teng Li 0009, Vikram K. Narayana, Tarek A. El-Ghazawi
ICPADS3
2011 GPU Resource Sharing and Virtualization on High Performance Computing Systems
abstract
Modern Graphic Processing Units (GPUs) are widely used as application accelerators in the High Performance Computing (HPC) field due to their massive floating-point computational capabilities and highly data-parallel computing architecture. Contemporary high performance computers equipped with co-processors such as GPUs primarily execute parallel applications using the Single Program Multiple Data (SPMD) model, which requires balanced computing resources of both microprocessor and co-processors to ensure full system utilization. While the inclusion of GPUs in HPC systems provides more computing resources and significant performance improvements, the asymmetrical distribution of the number of GPUs relative to the microprocessors can result in an underutilization of overall system computing resources. In this paper, we propose a GPU resource virtualization approach to allow underutilized microprocessors to share the GPUs. We analyze factors affecting the parallel execution performance on GPUs and conduct a theoretical performance estimation based on the most recent GPU architectures as well as the SPMD model. Then we present the implementation details of the virtualization infrastructure, followed by an experimental verification of the proposed concepts using an NVIDIA Fermi GPU computing node. The results demonstrate a considerable performance gain over the traditional SPMD execution without virtualization. Furthermore, the proposed solution enables full utilization of the asymmetrical system resources, through the sharing of the GPUs among microprocessors, while incurring low overheads due to the virtualization layer.
Teng Li 0009, Vikram K. Narayana, Esam El-Araby, Tarek A. El-Ghazawi
ICPP4
2011 New Hardware Architectures for Montgomery Modular Multiplication Algorithm
abstract
Montgomery modular multiplication is one of the fundamental operations used in cryptographic algorithms, such as RSA and Elliptic Curve Cryptosystems. At CHES 1999, Tenca and Koç proposed the Multiple-Word Radix-2 Montgomery Multiplication (MWR2MM) algorithm and introduced a now-classic architecture for implementing Montgomery multiplication in hardware. With parameters optimized for minimum latency, this architecture performs a single Montgomery multiplication in approximately 2n clock cycles, where n is the size of operands in bits. In this paper, we propose two new hardware architectures that are able to perform the same operation in approximately n clock cycles with almost the same clock period. These two architectures are based on precomputing partial results using two possible assumptions regarding the most significant bit of the previous word. These two architectures outperform the original architecture of Tenca and Koç in terms of the product latency times area by 23 and 50 percent, respectively, for several most common operand sizes used in cryptography. The architecture in radix-2 can be extended to the case of radix-4, while preserving a factor of two speedup over the corresponding radix-4 design by Tenca, Todorov, and Koç from CHES 2001. Our optimization has been verified by modeling it using Verilog-HDL, implementing it on Xilinx Virtex-II 6000 FPGA, and experimentally testing it using SRC-6 reconfigurable computer.
Miaoqing Huang, Kris Gaj, Tarek A. El-Ghazawi
IEEE Trans. Computers3
2011 A Framework for Evaluating High-Level Design Methodologies for High-Performance Reconfigurable Computers
abstract
High-performance reconfigurable computers have potential to provide substantial performance improvements over traditional supercomputers. Their acceptance, however, has been hindered by productivity challenges arising from increased design complexity, a wide array of custom design languages and tools, and often overblown sales literature. This paper presents a review and taxonomy of High-Level Languages (HLLs) and a framework for the comparative analysis of their features. It also introduces new metrics and a model based on computational effort. The proposed concepts are inspired by Netwon's equations of motion and the notion of work and power in an abstract multidimensional space of design specifications. The metrics are devised to highlight two aspects of the design process: the total time-to-solution and the efficient utilization of user and computing resources at discrete time steps along the development path. The study involves analytical and experimental evaluations demonstrating the applicability of the proposed model.
Esam El-Araby, Saumil G. Merchant, Tarek A. El-Ghazawi
IEEE Trans. Parallel Distributed Syst.3
2010 Reconfiguration and Communication-Aware Task Scheduling for High-Performance Reconfigurable Computing
abstract
High-performance reconfigurable computing involves acceleration of significant portions of an application using reconfigurable hardware. When the hardware tasks of an application cannot simultaneously fit in an FPGA, the task graph needs to be partitioned and scheduled into multiple FPGA configurations, in a way that minimizes the total execution time. This article proposes the Reduced Data Movement Scheduling (RDMS) algorithm that aims to improve the overall performance of hardware tasks by taking into account the reconfiguration time, data dependency between tasks, intertask communication as well as task resource utilization. The proposed algorithm uses the dynamic programming method. A mathematical analysis of the algorithm shows that the execution time would at most exceed the optimal solution by a factor of around 1.6, in the worst-case. Simulations on randomly generated task graphs indicate that RDMS algorithm can reduce interconfiguration communication time by 11% and 44% respectively, compared with two other approaches that consider data dependency and hardware resource utilization only. The practicality, as well as efficiency of the proposed algorithm over other approaches, is demonstrated by simulating a task graph from a real-life application - N-body simulation - along with constraints for bandwidth and FPGA parameters from existing high-performance reconfigurable computers. Experiments on SRC-6 are carried out to validate the approach.
Miaoqing Huang, Vikram K. Narayana, Harald Simmler, Olivier Serres, Tarek A. El-Ghazawi
ACM Trans. Reconfigurable Technol. Syst.5
2009 Efficient Mapping of Hardware Tasks on Reconfigurable Computers Using Libraries of Architecture Variants
abstract
Scheduling and partitioning of task graphs on reconfigurable hardware needs to be carefully carried out in order to achieve the best possible performance. In this paper, we demonstrate that a significant improvement to the total execution time is possible by incorporating a library of hardware task implementations, which contains multiple architectural variants for each hardware task reflecting tradeoffs between the resources utilization and the task execution throughput. We develop a genetic algorithm based mapping approach, which considers both task graph and target platform, and present results for an N-body simulation application using estimated numbers for resource utilization for the constituent tasks and based on actual architectural constraints from different reconfigurable platforms. The results demonstrate improvements of up to 85.3% in the execution time, compared to choosing a fixed implementation variant for each task while keeping a reasonable searching time.
Miaoqing Huang, Vikram K. Narayana, Tarek A. El-Ghazawi
FCCM3
2009 The Kamal Ewida Earth Observatory: A NATO Supported Real-time Remote Sensing Receiving Station being Established in Egypt with HPC-enabled Near-real-time Data Products for Mitigation of Environmental & Public Health Disasters
abstract
Establishment of the Kamal Ewida Earth Observatory (KEEO) has been funded by the North Atlantic Treaty Organization (NATO) Science for Peace Program. KEEO is a joint initiative of two of Egypt's largest and most venerable institutions of higher learning, Cairo University and Al Azhar University, both based in Cairo, Egypt, in collaboration with established environmental observatories in two NATO countries, Turkey and the USA. Specifically, the Egyptian partners, based in their Departments of Meteorology and Astronomy, Faculty of Science, at the two Egyptian Universities, are engaging in applications development, research and instructional collaboration with partnering resources from Bogaziçi University's Kandilli Observatory and Earthquake Research Institute (Istanbul, Turkey), with expertise in disaster mitigation, and Purdue University's Rosen Center for Advanced Computing's Purdue Terrestrial Observatory (West Lafayette, Indiana, USA), with expertise in real-time remote sensing and multi-disciplinary applications of satellite data. The KEEO project provides an interdisciplinary approach to effective disaster management and facilitates collaborative research and decision support, within the Egyptian context, for disaster mitigation.
Gilbert Rochon, Mohamed Magdy Abdel Wahab, Gamal Salah El Afandi, Gulay Altay, Okan K. Ersoy, Carol X. Song, Lan Zhao 0003, Larry L. Biehl, Belal Elleithy, Mohammed Shokr, Mohamed Mohamed 0002, Tarek A. El-Ghazawi, Darion Grant, Dev Niyogi
IGARSS (4)12
2009 RDMS: A hardware task scheduling algorithm for Reconfigurable Computing
abstract
Reconfigurable computers (RC) can provide significant performance improvement for domain applications. However, wide acceptance of today's RCs among domain scientist is hindered by the complexity of design tools and the required hardware design experience. Recent developments in HW/SW co-design methodologies for these systems provide the ease of use, but they are not comparable in performance to manual co-design. This paper aims at improving the overall performance of hardware tasks assigned to FPGA devices by minimizing both the communication overhead and configuration overhead, which are introduced by using FPGA devices. The proposed reduced data movement scheduling (RDMS) algorithm takes data dependency among tasks, hardware task resource utilization, and inter-task communication into account during the scheduling process and adopts a dynamic programming approach to reduce the communication between muP and FPGA co-processor and the number of FPGA configurations to a minimum. Compared to two other approaches that consider data dependency and hardware resource utilization only, RDMS algorithm can reduce inter-configuration communication time by 11% and 44% respectively based on simulation using randomly generated data flow graphs. The implementation of RDMS on a real-life application, N-body simulation, verifies the efficiency of RDMS algorithm against other approaches.
Miaoqing Huang, Harald Simmler, Olivier Serres, Tarek A. El-Ghazawi
IPDPS4
2009 Analytical modeling and evaluation of On-Chip Interconnects using Network Calculus
abstract
Network-on-Chip (NoC) has been proposed as an alternative to bus-based schemes to achieve high performance and scalability in System-on-Chip (SoC) design. Performance evaluation of On-Chip Interconnect (OCI) architectures is widely based on simulation which becomes computationally expensive, especially for large-scale NoCs. In this paper, a performance analysis model using Network Calculus is presented to characterize and evaluate the performance of NoC-based applications. The 2D Mesh on-chip interconnect is analyzed and main performance metrics such as end-to-end delay and buffer size requirements are computed and compared against the results produced by a discrete event simulator. The results shed more light on the potential of this analytical technique as a useful tool for NoC design and performance analysis.
Mohamed Bakhouya, Suboh A. Suboh, Jaafar Gaber, Tarek A. El-Ghazawi
NOCS4
2009 Virtual Configuration Management: A Technique for Partial Runtime Reconfiguration
abstract
Reconfigurable computers (RCs), built from configurable processors can offer high performance in a wide range of applications. However, due to the limited reconfigurable resources, not all needed functionalities can be implemented at the same time, and runtime reconfiguration becomes an appealing solution. This work proposes techniques suitable for multitasking applications as well as applications that can change the course of processing in a nondeterministic fashion. In order to exploit both spatial and temporal locality simultaneously, the proposed model groups hardware functions into configuration blocks of fixed size (pages), variable size (segments), or hybrid (paged segments). Multiple blocks can be configured on a chip simultaneously. Data mining techniques are used to group related functions into blocks (pages or segments) and temporal locality is exploited through block replacement techniques. Simulation, as well as emulation using the Cray XD1 reconfigurable high-performance computer was used in the experimental study. Results show a significant improvement in performance using the proposed techniques.
Mohamed Taher, Tarek A. El-Ghazawi
IEEE Trans. Computers2
2009 Exploiting Partial Runtime Reconfiguration for High-Performance Reconfigurable Computing
abstract
Runtime Reconfiguration (RTR) has been traditionally utilized as a means for exploiting the flexibility of High-Performance Reconfigurable Computers (HPRCs). However, the RTR feature comes with the cost of high configuration overhead which might negatively impact the overall performance. Currently, modern FPGAs have more advanced mechanisms for reducing the configuration overheads, particularly Partial Runtime Reconfiguration (PRTR). It has been perceived that PRTR on HPRC systems can be the trend for improving the performance. In this work, we will investigate the potential of PRTR on HPRC by formally analyzing the execution model and experimentally verifying our analytical findings by enabling PRTR for the first time, to the best of our knowledge, on one of the current HPRC systems, Cray XD1. Our approach is general and can be applied to any of the available HPRC systems. The paper will conclude with recommendations and conditions, based on our conceptual and experimental work, for the optimal utilization of PRTR as well as possible future usage in HPRC.
Esam El-Araby, Iván González 0004, Tarek A. El-Ghazawi
ACM Trans. Reconfigurable Technol. Syst.3
2008 Extreme parallel architectures for the masses
abstract
Multicore processors are now commodity items, and this has created an unprecedented buzz about exploiting parallelism to maximize performance. This is publicity has renewed interest in a long-standing problem: how much parallelism can we really exploit? Can extreme parallel computing be successfully delivered to the masses?
Tarek A. El-Ghazawi, Guy Lemieux
FPGA1
2008 Designing with extreme parallelism
abstract
Modern FPGAs can implement large, custom compute engines that are designed to exploit extreme amounts of parallel computation. Through parallelism, these systems achieve orders of magnitude higher performance than the fastest microprocessors. Building such custom compute engines with existing hardware design languages is too difficult and time-consuming. For this to become mainstream technology, the task of designing such parallel systems must be as simple as possible. Thus, high-level languages are needed which can specify a custom compute engine or be compiled to run on predesigned parallel systems. In this workshop, we will examine several approaches for specifying extremely parallel computations in high-level languages. These can be used to build parallel systems in FPGAs, or they can be used to specify parallel computations in other competing architectures. By examining several different approaches, one gains insight into the best approach for solving a given problem. Ideally, this will also inspire new approaches for designing with extreme parallelism
Guy Lemieux, Tarek A. El-Ghazawi
FPGA2
2008 Performance Evaluation of Clusters with ccNUMA Nodes - A Case Study
abstract
In the quest for higher performance and with the increasing availability of multicore chips, many systems are currently packing more processors per node. Adopting a ccNUMA node architecture in these cases has the promise of achieving a balance between cost and performance. In this paper, a 2312 Opteron cores system based on Sun Fire servers is considered as a case study to examine the performance issues associated with such architectures. In this work, we characterize the performance behavior of the system with focus on the node level using different configurations. It will be shown that the benefits from larger nodes can be severely limited due to many reasons. These reasons were isolated and the associated performance losses were assessed. The results revealed that such problems were mainly caused by topological imbalances, limitations of the used cache coherency protocol, operating system services distribution, and the lack of intelligent management of memory affinity.
Abdullah Kayi, Edward Kornkven, Tarek A. El-Ghazawi, Samy Al-Bahra, Gregory B. Newby
HPCC3
2008 Simulation and Evaluation of On-Chip Interconnect Architectures: 2D Mesh, Spidergon, and WK-Recursive Network
Suboh A. Suboh, Mohamed Bakhouya, Tarek A. El-Ghazawi
NOCS3
2008 Portable library development for reconfigurable computing systems: A case study
Proshanta Saha, Esam El-Araby, Miaoqing Huang, Mohamed Taher, Sergio López-Buedo, Tarek A. El-Ghazawi, Chang Shu 0003, Kris Gaj, Alan Michalski, Duncan A. Buell
Parallel Comput.6
2007 Applications of Heterogeneous Computing in Hardware/Software Co-Scheduling
abstract
Current work on automatic task partitioning and scheduling for reconfigurable computing (RC) systems strictly addresses the field programmable gate array (FPGA) hardware, and does not take advantage of the synergy between the microprocessor and the FPGA. Efforts on partitioning between the microprocessor and the FPGA are often times a manual and laborious effort as a formal methodology for automatic hardware-software partitioning for RC systems has not yet been established. Related fields such as heterogeneous computing (HC) and embedded computing (EC) have an extensive body of work for scheduling for heterogeneous processors. In this work, we adapt HC scheduling algorithms for RC systems, and show how simply adapting the algorithms alone is not sufficient to take advantage of the reconfigurable hardware. In many cases, the HC heuristics algorithms do not generate efficient schedules necessary to take advantage of the synergy between the microprocessor and the FPGA. We introduce new heuristic algorithms based on HC scheduling algorithms and show that they provide up to an order of magnitude improvement in execution time.
Proshanta Saha, Tarek A. El-Ghazawi
AICCSA2
2007 Software/Hardware Co-Scheduling for Reconfigurable Computing Systems
abstract
A formal methodology for automatic hardware-software partitioning and co-scheduling between the P and the FPGA has not yet been established. Current work in automatic task partitioning and scheduling for the reconfigurable systems strictly addresses the FPGA hardware, and does not take advantage of the synergy between the microprocessor and the FPGA. In this work, we consider the problem of co-scheduling task graphs on reconfigurable systems. The target systems have an execution model which allows any subtask that can run on the FPGA to also run on the microprocessor, and allows reconfigurability of the FPGA (subject to area, performance, resource, and timing constraints). In this paper, we introduce a new heuristic algorithm for such hardware/software co-scheduling, ReCoS. It will be shown that the proposed algorithm provides up to an order of magnitude improvement in scheduling and execution times when compared with hardware/software co-schedulers found in the embedded systems area, after adapting them for reconfigurable computing.
Proshanta Saha, Tarek A. El-Ghazawi
FCCM2
2007 Bringing High-Performance Reconfigurable Computing to Exact Computations
abstract
Numerical non-robustness is a recurring phenomenon in scientific computing. It is primarily caused by numerical errors arising because of fixed-precision arithmetic in integer and/or floating-point computations. Exact computation, based on arbitrary-precision arithmetic, has been developed over the last decade as an emerging numerical computation paradigm in response to this problem of numerical non-robustness. Exact arithmetic, specifically arbitrary-precision arithmetic, has been traditionally implemented using efficient software libraries such as GNU Multi-Precision (GMP). However, this results in a slower arithmetic performance when compared to fixed-precision arithmetic. In this paper we present a first effort, to the best of our knowledge, of reconfigurable hardware support for arbitrary-precision arithmetic. The proposed hardware architectures are based on virtual convolution sche1duling which is derived from a formal representation of the problem. Targeting high performance and efficiency, dynamic (non-linear) pipelines techniques were exploited to eliminate the effects of deeply-pipelined operators. Referenced to GMP, our experiments showed promising results.
Esam El-Araby, Iván González 0004, Tarek A. El-Ghazawi
FPL3
2007 Productivity of High-Level Languages on Reconfigurable Computers: An HPC Perspective
abstract
Productivity on high-performance reconfigurable computers (HPRCs) is becoming a concern given the complexity of today's applications and development flows. Furthermore, the plethora of options from which application developers need to select their development environments has recently become another productivity obstacle. High-level languages (HLLs) for developing reconfigurable computing applications trade performance with ease-of-use. However, it is hard to know in a general sense how much performance one is giving up and how much ease-of-use he/she is gaining. More importantly, given the lack of standards and the uncertainty generated by sales literature, it is very hard to know the real differences that exist among different high-level programming paradigms. In order to do so, one needs a classification of HLLs programming models from a general high-performance computing (HPC) perspective. In this work, we consider a number of representative high-level tools that were selected to represent imperative programming, functional programming and graphical programming, and thereby demonstrate the applicability of our methodology. It will be shown that in spite of the disparity in concepts behind those tools, our methodology will be able to uncover the basic differences among them and assess their comparative productivity in terms of performance, and ease-of-use.
Esam El-Araby, Preetham Nosum, Tarek A. El-Ghazawi
FPT3
2007 A Portable Memory Access Framework on Reconfigurable Computers
abstract
Current reconfigurable computers (RCs) do not share a unified architectural model, which presents a challenge to any developer who intends to port hardware designs across different RC platforms. In this paper, we propose a portable memory access framework that gives the user a unified memory view combining the host memory and the local memory of FPGA. Three memory access modes are provided, and the hardware cost and performance impact have been measured on three major RCs: SRC-6, SGI RC-100 and Cray XD1. Under current implementation, the penalty to hardware resource utilization and performance of applying this framework is reduced to minimum.
Miaoqing Huang, Iván González 0004, Tarek A. El-Ghazawi
FPT3
2007 Towards a Complexity Model for Design and Analysis of PGAS-Based Algorithms
Mohamed Bakhouya, Jaafar Gaber, Tarek A. El-Ghazawi
HPCC3
2007 Experimental Evaluation of Emerging Multi-core Architectures
abstract
The trend of increasing speed and complexity in the single-core processor as stated in the Moore's law is facing practical challenges. As a result, the multi-core processor architecture has emerged as the dominant architecture for both desktop and high-performance systems. Multi-core systems introduce many challenges that need to be addressed to achieve the best performance. Therefore, a new set of benchmarking techniques to study the impacts of the multi-core technologies is necessary. In this paper, multi-core specific performance metrics for cache coherency and memory bandwidth/latency/contention are investigated. This study also proposes a new benchmarking suite which includes cases extended from the high performance computing challenge (HPCC) benchmark suite. Performance results are measured on a Sun Fire T1000 server with six cores and an AMD Opteron dual core system. Experimental analysis and observations in this paper provide for a better understanding of the emerging multi-core architectures.
Abdullah Kayi, Yiyi Yao, Tarek A. El-Ghazawi, Gregory B. Newby
IPDPS3
2007 A Methodology for Automating Co-Scheduling for Reconfigurable Computing Systems
abstract
A formal methodology for automatic hardware-software partitioning and co-scheduling between the muP and the FPGA has not yet been established. Current work in automatic task partitioning and scheduling for reconfigurable systems strictly addresses the FPGA hardware, and does not take advantage of the synergy between the microprocessor and the FPGA. In this work, we consider the problem of co-scheduling task graphs on reconfigurable systems. The target systems have an execution model which allows any subtask that can run on the FPGA to also run on the microprocessor, and allows reconfigurability of the FPGA (subject to area, performance, resource, and timing constraints). In this paper, we introduce a methodology for automatic co- scheduling using a proposed heuristic algorithm for hardware/software co-scheduling, ReCoS. It will be shown that the proposed algorithm provides up to an order of magnitude improvement in scheduling and execution times when compared with hardware/software co-schedulers found in related fields such as embedded systems, heterogeneous systems, and reconfigurable hardware systems.
Proshanta Saha, Tarek A. El-Ghazawi
MEMOCODE2
2007 Hierarchical PCA Techniques for Fusing Spatial and Spectral Observations With Application to MISR and Monitoring Dust Storms
abstract
In this letter, we propose hierarchical principal component analysis (HPCA) techniques for fusing spatial and spectral data, and compare them to direct principal component analysis (DPCA) over Multiangle Imaging SpectroRadiometer (MISR) data. It is shown that the proposed methods are significantly faster than DPCA. In case of DPCA, we merge the 20 different images resulting from the four spectral bands over the nadir and the four forward angles. In the hierarchical case, we first merge the information from the four spectral camera bands; then, we integrate the spatial information from the five cameras in the second step (or vice versa) by applying principal component analysis (PCA) twice. The classification results show that fused data using HPCA compare favorably to DPCA or to classification using the original data. This is because applying PCA to one particular data domain (e.g., spectral data followed by spatial data or vice versa) tends to better remove redundancies and enhance features within that domain. In addition, classification through hierarchical data fusion results in computational savings over the other methods.
Hesham Mohamed El-Askary, Tarek A. El-Ghazawi, Menas Kafatos, Jacqueline LeMoigne-Stewart
IEEE Geosci. Remote. Sens. Lett.3
2006 A Segmentation Model for Partial Run-Time Reconfiguration
abstract
Reconfigurable computing systems have been gaining rising attention. Such systems adapt the overall system to the underlying applications at run-time. However, due to the limited reconfigurable resources, not all needed functionalities can be implemented in the same time. Previous work has considered swapping hardware functions on either a function-by-function basis or by reconfiguring the whole chip. In our previous work we have proposed a configuration management technique based on grouping related functions into fixed size blocks (pages). Pages are swapped in and out as necessary during application execution. However, paging can introduce physical artificial constraints on the grouping decision. In this work, we propose a more general virtual-memory-like technique. This technique discovers related functions and groups them into variable size blocks (segments). This, in addition to block replacement strategies can exploit both spatial and temporal processing locality simultaneously. Simulations, as well as emulation have been performed using the Cray XD1 reconfigurable computer. Results have shown that the proposed model can provide several folds of speed-up over previous techniques
Mohamed Taher, Tarek A. El-Ghazawi
FPL2
2006 Exploiting processing locality through paging configurations in multitasked reconfigurable systems
abstract
FPGA chips in reconfigurable computer systems are used as malleable coprocessors where components of a hardware library of functions can be configured as needed. As the number of hardware functions to be configured typically exceeds the underlying chip area during the execution of an application, previous efforts have introduced configuration caching. Those efforts, however, have focused on two run-time-reconfiguration scenarios, which consider a single application running on the reconfigurable system. In the full reconfiguration scenario, functions of an application are arranged into blocks each of which has enough functions to fill the entire chip. The blocks are configured in a deterministic sequence needed by the application based on the a priori knowledge about the application. In the partial reconfiguration scenario, each function is configured or replaced on a function-by-function basis, based on the application needs. In the former technique, spatial processing locality is well exploited. In the latter, only temporal processing locality is exploited. In this work, we propose a technique suitable for multitasking and for cases of single applications that can change the course of processing in a non-deterministic fashion based on data. In order to exploit processing locality, both spatial and temporal simultaneously, the proposed model groups hardware functions into hardware configuration blocks (pages) of fixed size, where multiple pages can be configured on a chip simultaneously. By grouping only related functions that are typically requested together, processing spatial locality can be exploited. Temporal locality is exploited through page replacement techniques. Data mining techniques were used to group related functions into pages. Standard, replacement algorithms as those found in caching were considered. Simulations, as well as emulation using the Cray XD1 reconfigurable high-performance computer were used in the experimental study. The results show a significant improvement in performance using the proposed paging technique.
Mohamed Taher, Tarek A. El-Ghazawi
IPDPS2
2006 M03 - Reconfigurable supercomputing
abstract
The synergistic advances in high-performance computing and reconfigurable computing, based on field programmable gate arrays (FPGAs), has resulted in hybrid parallel systems of microprocessors and FPGAs. Such systems support both fine-grain and coarse-grain parallelism, and can dynamically tune their architecture to fit various applications. Programming these systems can be quite challenging as programming of FPGA devices can involve hardware design. This tutorial will introduce the field of reconfigurable supercomputing and its advances in systems, programming, applications and tools. Reconfigurable system developments at SRC, Cray, SGI, and Star Bridge will be highlighted, and case studies including full application developments will be presented along with the live demonstrations. This tutorial will be the first to show scalability studies for real-life applications over entire HPRC systems. This will reveal the tremendous promise held by this class of architectures in performance, power and cost improvements. Challenges that remain will be also discussed.
Tarek A. El-Ghazawi, Duncan A. Buell, Volodymyr V. Kindratenko, Kris Gaj
SC1
2006 Reconfigurable supercomputing - Is high-performance reconfigurable computing the next supercomputing paradigm?
abstract
High-Performance Reconfigurable Computers (HPRCs) based on the combination of conventional processors and FPGAs have been gaining attention in the past few years. Their benefits were particularly harnessed in compute-intensive integer applications. However, there has been doubt that the same benefits can be attained for general scientific applications. Fortunately, the trend in reconfigurable chip sizes and diversity of resources may be relieving some of those concerns. Yet, with the hardware reconfigurability, it is feared that domain scientists have to learn how to design hardware if they were to use such machines effectively. In order to address the overarching question, this panel will address the following questions: Can FPGAs deliver order-of-magnitude performance gains in scientific floating-point applications in the foreseeable future? Can programming HPRCs programmability become similar to that of HPCs in its level of difficulty? What are the major developments in the industry or the community that make all this possible?
Tarek A. El-Ghazawi, Dave Bennett, Daniel S. Poznanovic, Allan Cantle, Keith D. Underwood, Rob Pennington, Duncan A. Buell, Alan D. George, Volodymyr V. Kindratenko
SC1
2006 UPC - UPC: unified parallel C
abstract
UPC extends ISO C into a Partioned Global Address Space (PGAS) programming language. UPC allows programmers to exploit data locality and parallelism in their applications, while maintaining ease of use. UPC is running ubiquitously across nearly all HPC platforms and has been gaining rising support from the community. UPC is relatively very easy to use for irregular access patterns which can enable many new applications that are hard to express in other paradigms. In this BoF, the UPC consortium will share with the community the progress made in applications, specifications, tools, and implementations of UPC. Future plans will be also presented. The First half of the BOF will follow a panel format, where the members will represent the UPC consortium and will speak to different aspects of the UPC developments. The second half will be a question and answer session to promote the exchange of ideas.
Tarek A. El-Ghazawi, Lauren Smith
SC1
2006 Benchmarking parallel compilers: A UPC case study
Tarek A. El-Ghazawi, François Cantonnet, Yiyi Yao, Smita Annareddy, Ahmed S. Mohamed
Future Gener. Comput. Syst.1
2005 Reconfigurable computers: an empirical analysis (abstract only)
abstract
Reconfigurable Computers are parallel systems that are designed around multiple general-purpose processors and multiple field programmable gate array (FPGA) chips. These systems can leverage the synergism between conventional processors and FPGAs to provide low-level hardware functionality at the same level of programmability as general-purpose computers. In this work we conduct an experimental study using one of the state-of-the-art reconfigurable computers and a representative set of applications to assess the field, uncover the challenges, propose solutions, and conceive a realistic evolution path. We consider issues of concern including performance/cost. We also consider productivity in the sense of development, compiling, running, and system reliability. It will be shown that for some applications, the performance/cost can be orders of magnitude better than conventional computers. It will be also shown that programming such machines may still require some hardware knowledge, similar to hardware knowledge computer programmers must acquire to write scalable programs.
Tarek A. El-Ghazawi, Kris Gaj, Nikitas A. Alexandridis, Allen Michalski, Osman Devrim Fidanci, Mohamed Taher, Esam El-Araby, Esmail Chitalwala, Proshanta Saha
FPGA1
2005 Image processing library for reconfigurable computers (abstract only)
abstract
Reconfigurable Computers (RCs) are parallel systems that are designed around multiple general-purpose processors and multiple field programmable gate array (FPGA) chips. These systems can leverage the synergism between conventional processors and FPGAs to provide low-level hardware functionality at the same level of programmability as general-purpose computers. RCs have proposed very high processing capabilities for computationally intensive applications such as Image Processing. This is due to the inherently parallel operation paradigm of the FPGA hardware.In this paper we present the design and implementation of image processing kernels for RCs. This library of kernels have been tested and verified for performance on one of the state-of-the-art reconfigurable computers, SRC-6E. This paper shows that RCs are between 8 to 400 times faster than comparable Pentiums for image based tasks.
Mohamed Taher, Esam El-Araby, Tarek A. El-Ghazawi, Kris Gaj
FPGA3
2005 A System-Level Design Methodology for Reconfigurable Computing Applications
Esam El-Araby, Tarek A. El-Ghazawi, Kris Gaj
FPT2
2005 Prototyping Automatic Cloud Cover Assessment (ACCA) Algorithm for Remote Sensing On-Board Processing on a Reconfigurable Computer
Esam El-Araby, Mohamed Taher, Tarek A. El-Ghazawi, Jacqueline LeMoigne-Stewart
FPT3
2005 Performance of Sorting Algorithms on the SRC 6 Reconfigurable Computer
John Harkins, Tarek A. El-Ghazawi, Esam El-Araby, Miaoqing Huang
FPT2
2005 Low Latency Elliptic Curve Cryptography Accelerators for NIST Curves Over Binary Fields
Chang Shu 0003, Kris Gaj, Tarek A. El-Ghazawi
FPT3
2005 An efficient implementation of on-board cloud detection on a reconfigurable computer
abstract
The presence of cloud contamination can hinder the use of satellite data, and this requires a cloud detection process to mask out cloudy pixels from further processing. The trend for remote sensing satellite missions has always been towards smaller size, lower cost, more flexibility, and higher computational power. Reconfigurable Computers (RCs) combine the flexibility of traditional microprocessors with the power of Field Programmable Gate Arrays (FPGAs). Therefore, RCs are a promising candidate for on-board preprocessing. This paper presents the design and implementation of an RC-based real-time cloud detection system. We investigate the potential of using RCs for on-board preprocessing by prototyping the Landsat 7 ETM+ ACCA algorithm on one of the state-of-the art reconfigurable platforms, SRC-6E. Although a reasonable amount of investigations of the ACCA cloud detection algorithm using FPGAs has been reported in the literature, very few details/results were provided and/or limited contributions were accomplished. Our work has been proven to provide higher performance and higher detection accuracy.
Esam El-Araby, Mohamed Taher, Tarek A. El-Ghazawi, Jacqueline LeMoigne-Stewart
IGARSS3
2005 Enhancing dust storm detection using PCA based data fusion
abstract
Principal Component Analysis (PCA) has been widely used as a data reduction technique to overcome the curse of dimensionality. In this research we show a different use for PCA technique as a tool for data fusion. PCA as a data fusion technique is performed over the Multiangle Imaging Spectroradiometer (MISR) data, studying dust storms to better serve their identification. The multi-angle viewing capability of MISR is used to enhance our understanding of the Earth's environment that includes climate particularly of atmosphere and of land surfaces. In this research the multi angle MISR images clearly show a dust storm over the Liaoning region of China as well as parts of northern and western Korea on April 8, 2002. PCA is used to combine the obtained information from the different angle views and frequency bands of MISR datasets. Performing K-means clustering on the original and the assimilated products apply a quantitative measure that is introduced. Upon classifying the first 4 principal components (PCs) having 95% of the information content similar results were obtained as compared to the classification using original datasets.
Hesham Mohamed El-Askary, Tarek A. El-Ghazawi, Menas Kafatos, Jacqueline LeMoigne-Stewart
IGARSS3
2005 An evaluation of global address space languages: co-array fortran and unified parallel C
abstract
Co-array Fortran (CAF) and Unified Parallel C (UPC) are two emerging languages for single-program, multiple-data global address space programming. These languages boost programmer productivity by providing shared variables for inter-process communication instead of message passing. However, the performance of these emerging languages still has room for improvement. In this paper, we study the performance of variants of the NAS MG, CG, SP, and BT benchmarks on several modern architectures to identify challenges that must be met to deliver top performance. We compare CAF and UPC variants of these programs with the original Fortran+MPI code. Today, CAF and UPC programs deliver scalable performance on clusters only when written to use bulk communication. However, our experiments uncovered some significant performance bottlenecks of UPC codes on all platforms. We account for the root causes limiting UPC performance such as the synchronization model, the communication efficiency of strided data, and source-to-source translation issues. We show that they can be remedied with language extensions, new synchronization constructs, and, finally, adequate optimizations by the back-end C compilers.
Cristian Coarfa, Yuri Dotsenko, John M. Mellor-Crummey, François Cantonnet, Tarek A. El-Ghazawi, Ashrujit Mohanti, Yiyi Yao, Daniel G. Chavarría-Miranda
PPoPP5
2004 Implementation of elliptic curve cryptosystems over GF(2n) in optimal normal basis on a reconfigurable computer
abstract
During the last few years, a considerable effort has been devoted to the development of reconfigurable computers, machines that are based on the close interoperation of traditional microprocessors and Field Programmable Gate Arrays. Several prototype machines of this type have been designed, and demonstrated significant speed-ups compared to conventional workstations for computationally intensive problems, such as codebreaking. In this paper, we demonstrate an efficient implementation of Elliptic Curve scalar multiplication over GF(2 n ) in Optimal Normal Basis, using one of the leading reconfigurable computers available on the market, SRC-6E. We show how the hardware architecture and programming model of this reconfigurable computer has influenced the choice of the optimum program partitioning scheme. The detailed analysis of the control, data transfer, and reconfiguration overheads is given in the paper. The end-to-end speed-ups in the range from 895 to 1300 compared to the microprocessor implementation are demonstrated depending on the chosen partitioning scheme.
Sashisu Bajracharya, Chang Shu 0003, Kris Gaj, Tarek A. El-Ghazawi
FPGA4
2004 Implementation of Elliptic Curve Cryptosystems over GF(2n) in Optimal Normal Basis on a Reconfigurable Computer
Sashisu Bajracharya, Chang Shu 0003, Kris Gaj, Tarek A. El-Ghazawi
FPL4
2004 Run-Time Reconfiguration Management for Adaptive High-Performance Computing Systems
Mohamed Taher, Tarek A. El-Ghazawi
FPL2
2004 Reconfigurable hardware implementation of mesh routing in number field sieve factorization
abstract
Factorization of large numbers has been a constant source of interest in cryptanalysis. The fastest known algorithm for factoring large numbers is the number field sieve (NFS). The two most time consuming phases of NFS are sieving and matrix step. We propose an efficient way of implementing the matrix step in reconfigurable hardware. Our solution is based on the mesh-routing method proposed by Lenstra et al. We determine the practical size of a partial mesh that can fit in one FFGA device, Xilinx Virtex II XC2V6000. We further extrapolate the computation time for the case of a square systolic array of FFGAs for 512-bit and 1024-bit numbers' factorization. We demonstrate that for practical sizes of numbers used in cryptography, 1024 bits, the matrix step of factorization can be performed using 1024 Virtex II FFGAs in less than 40 days.
Sashisu Bajracharya, Deapesh Misra, Kris Gaj, Tarek A. El-Ghazawi
FPT4
2004 Effective system and performance benchmarking for reconfigurable computers
abstract
Applications running on a reconfigurable computer can be divided into two major categories: computationally intensive and input/output intensive. In the first case, the input and output are limited, and therefore the performance of the reconfigurable computer depends primarily on the power of the FPGAs, and the capability to exploit parallelism available in a given application. In the second case, the execution time is dominated by input/output, and therefore, an application cannot process data faster than the speed of its slowest input/output channel. The focus of This work is on developing micro-benchmarks to characterize the behavior of various communication channels within reconfigurable computers for the second class of applications. The paper defines a system of 'paper and pencil' micro-benchmarks for the measurement of maximum throughput and minimum latency in the communication between various components of a generic reconfigurable system. The results help to dynamically characterize a reconfigurable machine and the SRC 6E reconfigurable computer is used as a test case to validate the proposed model.
Esmail Chitalwala, Tarek A. El-Ghazawi, Kris Gaj, Nikitas A. Alexandridis, Daniel S. Poznanovic
FPT2
2004 Wavelet spectral dimension reduction of hyperspectral imagery on a reconfigurable computer
abstract
Hyperspectral imagery, by definition, provides valuable remote sensing observations at hundreds of frequency bands. Conventional image classification (interpretation) methods may not be used without dimension reduction preprocessing. Automatic wavelet reduction has been proven to yield better or comparable classification accuracy, while achieving substantial computational savings. However, the large hyperspectral data volumes remain to present a challenge for traditional processing techniques. Reconfigurable computers (RCs) can leverage the synergism between conventional processors and FPGAs to provide low-level hardware functionality at the same level of programmability as general-purpose computers. We investigate the potential of using RCs for on-board, i.e. aboard airborne/spaceborne carriers, preprocessing of hyperspectral imagery by prototyping for the first time the automatic wavelet dimension reduction algorithm. Our investigation exploits the fine and coarse grain parallelism provided by the RCs and has been experimentally verified on one of the state-of the art reconfigurable platforms, SRC-6E. An order of magnitude speedup over traditional processing techniques has been reported.
Esam El-Araby, Tarek A. El-Ghazawi, Jacqueline LeMoigne-Stewart, Kris Gaj
FPT2
2004 Wavelet dimension reduction of AIRS infrared (IR) hyperspectral data
abstract
Recently developed hyperspectral sensors provide much richer information than comparable multispectral sensors. However traditional methods that have been designed for multispectral data are not easily adaptable to hyperspectral data. One way to approach this problem is to perform dimension reduction as pre-processing, i.e. to apply a transformation that brings data from a high order dimension to a low order dimension. Wavelet spectral analysis of hyperspectral images has been recently proposed as a method for dimension reduction and, when tested on the classification of AVIRIS data, has shown promising results over the traditional principal component analysis (PCA) technique. We propose to extend and apply the wavelet analysis reduction method to the Atmospheric Infrared Sounder (AIRS) instrument data, designed to measure the Earth's atmospheric water vapor and temperature profiles on a global scale. With more than 2,000 channels, the AIRS infrared data represent a good candidate for dimension reduction, and especially wavelet reduction, due to its computational efficiency and the large data sizes involved.
Jacqueline LeMoigne-Stewart, Joanna Joiner, Tarek A. El-Ghazawi, François Cantonnet
IGARSS4
2004 Productivity Analysis of the UPC Language
abstract
Summary form only given. Parallel programming paradigms, over the past decade, have focused on how to harness the computational power of contemporary parallel machines. Ease of use and code development productivity, has been a secondary goal. Recently, however, there has been a growing interest in understanding the code development productivity issues and their implications for the overall time-to-solution. Unified Parallel C (UPC) is a recently developed language which has been gaining rising attention. UPC holds the promise of leveraging the ease of use of the shared memory model and the performance benefit of locality exploitation. The performance potential for UPC has been extensively studied in recent research efforts. The aim of this study, however, is to examine the impact of UPC on programmer productivity. We propose several productivity metrics and consider a wide array of high performance applications. Further, we compare UPC to the most widely used parallel programming paradigm, MPI. The results show that UPC compares favorably with MPI in programmers productivity.
François Cantonnet, Yiyi Yao, Mohamed Zahran 0001, Tarek A. El-Ghazawi
IPDPS4
2004 System-Level Parallelism and Throughput Optimization in Designing Reconfigurable Computing Applications
abstract
Summary form only given. Reconfigurable computers (RCs) can leverage the synergism between conventional processors and FPGAs to provide low-level hardware functionality at the same level of programmability as general-purpose computers. In a large class of applications, the total I/O time is comparable or even greater than the computations time. As a result, the rate of the DMA transfer between the microprocessor memory and the on-board memory of the FPGA-based processor becomes the performance bottleneck. We perform a theoretical and experimental study of this specific performance limitation. The mathematical formulation of the problem has been experimentally verified on the state-of-the art reconfigurable platform, SRC-6E. We demonstrate and quantify the possible solution to this problem that exploits the system-level parallelism within reconfigurable machines.
Esam El-Araby, Mohamed Taher, Kris Gaj, Tarek A. El-Ghazawi, David Caliga, Nikitas A. Alexandridis
IPDPS4
2004 A performance study of job management systems
abstract
Abstract Job Management Systems (JMSs) efficiently schedule and monitor jobs in parallel and distributed computing environments. Therefore, they are critical for improving the utilization of expensive resources in high‐performance computing systems and centers, and an important component of Grid software infrastructure. With many JMSs available commercially and in the public domain, it is difficult to choose an optimum JMS for a given computing environment. In this paper, we present the results of the first empirical study of JMSs reported in the literature. Four commonly used systems, LSF, PBS Pro, Sun Grid Engine/CODINE, and Condor were considered. The study has revealed important strengths and weaknesses of these JMSs under different operational conditions. For example, LSF was shown to exhibit excellent throughput for a wide range of job types and submission rates. Alternatively, CODINE appeared to outperform other systems in terms of the average turn‐around time for small jobs, and PBS appeared to excel in terms of turn‐around time for relatively larger jobs. Copyright © 2004 John Wiley & Sons, Ltd.
Tarek A. El-Ghazawi, Kris Gaj, Nikitas A. Alexandridis, Frederic Vroman, Nguyen Nguyen 0003, Jacek R. Radzikowski, Preeyapong Samipagdi, Suboh A. Suboh
Concurr. Pract. Exp.1
2003 An Implementation Comparison of an IDEA Encryption Cryptosystem on Two General-Purpose Reconfigurable Computers
Allen Michalski, Kris Gaj, Tarek A. El-Ghazawi
FPL3
2003 Exploiting system-level parallelism in the application development on a reconfigurable computer
abstract
Reconfigurable Computers (RCs) can leverage the synergism between conventional processors and FPGAs to provide low-level hardware functionality at the same level of programmability as general-purpose computers. In a large class of applications, the total I/O time is comparable or even greater than the computations time. As a result, the rate of the DMA transfer between the microprocessor memory and the on-board memory becomes the performance bottleneck even on RCs. In this paper, we perform a theoretical and experimental study of this specific performance limitation for the state-of-the art reconfigurable platform, SRC-6E. We demonstrate and quantify the possible solution to this problem that exploits the system-level parallelism within the reconfigurable machine.
Esam El-Araby, Mohamed Taher, Kris Gaj, Tarek A. El-Ghazawi, David Caliga, Nikitas A. Alexandridis
FPT4
2003 Implementation of Elliptic Curve Cryptosystems on a reconfigurable computer
abstract
During the last few years, a considerable effort has been devoted to the development of reconfigurable computers, machines that are based on the close interoperation of traditional microprocessors and Field Programmable Gate Arrays (FPGAs). Several prototype machines of this type have been designed, and demonstrated significant speedups compared to conventional workstations for computationally intensive problems, such as codebreaking. Nevertheless, the efficient use and programming of such machines is still an unresolved problem. In this paper, we demonstrate an efficient implementation of an Elliptic Curve scalar multiplication over GF(2/sup m/), using one of the leading reconfigurable computers available on the market, SRC-6E. We show how the hardware architecture and programming model of this reconfigurable computer has influenced the choice of the algorithm partitioning strategy for this application. A detailed analysis of the control, data transfer, and reconfiguration overheads is given in the paper, together with the performance comparison of our implementation against an optimized microprocessor implementation.
Nghi Nguyen, Kris Gaj, David Caliga, Tarek A. El-Ghazawi
FPT4
2003 Introducing new approaches for dust storms detection using remote sensing technology
abstract
Dust storms present environmental risks and affect the climate. They have worsened in the Mediterranean and East Asia regions over the last decade due to massive deforestation and increased droughts. Storms can travel over large parts of the Earth, in Asia, Africa, even affecting North America and Europe. Moreover dust storms are related to precipitation, soil moisture, land use/land cover practices, and other human activities. This work is a continuation of previous research in which we analyzed several remote sensing instruments capabilities in monitoring dust storms. We introduce the usage of the Multi-angle Imaging SpectroRadiometer (MISR) and TRMM Microwave Imager (TMI) as an optical and microwave combination in enhancing dust storm detection.
Hesham Mohamed El-Askary, Menas Kafatos, Tarek A. El-Ghazawi
IGARSS4
2003 A self-stabilizing distributed algorithm for spanning tree construction in wireless ad hoc networks
Hichem Baala, Olivier Flauzac, Jaafar Gaber, Marc Bui, Tarek A. El-Ghazawi
J. Parallel Distributed Comput.5
2003 A multisensor approach to dust storm monitoring over the Nile Delta
abstract
This work analyzes several remote sensing instrument capabilities in monitoring dust storms. Multisensor data analysis is carried out to study the behavior of dust particles at different wavelengths. A technique based on a combination of optical and microwave sensing of dust storms, using the Moderate Resolution Imaging Spectrometer (MODIS) and the Tropical Rainfall Measuring Mission (TRMM) Microwave Imager (TMI) respectively, is found to be particularly useful.
Hesham Mohamed El-Askary, Sudipta Sarkar, Menas Kafatos, Tarek A. El-Ghazawi
IEEE Trans. Geosci. Remote. Sens.4
2003 Automatic reduction of hyperspectral imagery using wavelet spectral analysis
abstract
Hyperspectral imagery provides richer information about materials than multispectral imagery. The new larger data volumes from hyperspectral sensors present a challenge for traditional processing techniques. For example, the identification of each ground surface pixel by its corresponding spectral signature is still difficult because of the immense volume of data. Conventional classification methods may not be used without dimension reduction preprocessing. This is due to the curse of dimensionality, which refers to the fact that the sample size needed to estimate a function of several variables to a given degree of accuracy grows exponentially with the number of variables. Principal component analysis (PCA) has been the technique of choice for dimension reduction. However, PCA is computationally expensive and does not eliminate anomalies that can be seen at one arbitrary band. Spectral data reduction using automatic wavelet decomposition could be useful. This is because it preserves the distinctions among spectral signatures. It is also computed in automatic fashion and can filter data anomalies. This is due to the intrinsic properties of wavelet transforms that preserves high- and low-frequency features, therefore preserving peaks and valleys found in typical spectra. Compared to PCA, for the same level of data reduction, we show that automatic wavelet reduction yields better or comparable classification accuracy for hyperspectral data, while achieving substantial computational savings.
Sinthop Kaewpijit, Jacqueline LeMoigne-Stewart, Tarek A. El-Ghazawi
IEEE Trans. Geosci. Remote. Sens.3
2002 A wavelet-based PCA reduction for hyperspectral imagery
abstract
Hyperspectral Imagery can provide very rich information on land cover classes. However, it also presents many challenges in data analysis and interpretation, due to the large amount of data collected. For example, conventional methods for land use and land cover classifications may not be applicable, due to "the curse of dimensionality." Therefore, these conventional methods may need a preprocessing step to transform high dimensional data to low dimensional data, by eliminating data redundancy. Due to its conceptual simplicity, principal component analysis (PCA) has been widely used for decades to reduce dimensionality. It is a useful technique if the spectral class structure of the transformed data is such that it is distributed along the first few axes. Otherwise, the transformed data may be similar to the original data. In such a case, the wavelet decomposition technique might be a better approach. Wavelet decomposition can reduce hyperspectral data in the spectral domain for each pixel. This will not only reduce the data volume, but will also preserve the distinction among spectral signatures that is useful for most pixel-based classifiers. This characteristic is related to the intrinsic property of wavelet transforms that preserve high- and low-frequency features during the signal decomposition, and therefore preserve peaks and valleys found in typical spectra. In general, most classification errors occur at the boundary between classes. Since wavelet decomposition is applied to each local pixel, a wavelet-based reduction might not well differentiate classes among neighborhood pixels in the spatial domain. PCA, however, can provide more local spatial information among neighborhood class pixels than wavelet.
Sinthop Kaewpijit, Jacqueline LeMoigne-Stewart, Tarek A. El-Ghazawi
IGARSS3
2002 UPC performance and potential: a NPB experimental study
abstract
UPC, or Unified Parallel C, is a parallel extension of ANSI C. UPC follows a distributed shared memory programming model aimed at leveraging the ease of programming of the shared memory paradigm, while enabling the exploitation of data locality. UPC incorporates constructs that allow placing data near the threads that manipulate them to minimize remote accesses. This paper gives an overview of the concepts and features of UPC and establishes, through extensive performance measurements of NPB workloads, the viability of the UPC programming language compared to the other popular paradigms. Further, through performance measurements we identify the challenges, the remaining steps and the priorities for UPC. It will be shown that with proper hand tuning and optimized collective operations libraries, UPC performance will be comparable to that of MPI. Furthermore, by incorporating such improvements into automatic compiler optimizations, UPC will compare quite favorably to message passing in ease of programming.
Tarek A. El-Ghazawi, François Cantonnet
SC1
2001 Parallel and Adaptive Reduction of Hyperspectral Data to Intrinsic Dimensionality
abstract
Recent advances in sensor technology have led to the development of hyperspectral sensors capable of collecting remote sensing imagery at several hundred bands over the spectrum. While these developments hold great promise for Earth science, they create new processing challenges. Therefore, processing hyperspectral data using new efficient techniques is a must.One such approach of effectively processing hyperspectral data is through dimension reduction prior to the use of data in applications such as land use/land cover classifications. Principal Component Analysis (PCA) is perhaps the most popular dimension reduction technique. The outcome is a number of principal components (PCs) in a descending order of information content. It is often the case that only a small number of these components contain the effective information needed. In this paper, we propose an adaptive parallel technique for determining the effective dimensionality of hyperspectral data on computer clusters. Based on a user-specified desired level of information content, the method selects adaptively the faster technique for solving the eigenproblem and computes only the needed components for that level of information.
Tarek A. El-Ghazawi, Sinthop Kaewpijit, Jacqueline LeMoigne-Stewart
CLUSTER1
2001 UPC Benchmarking Issues
abstract
UPC, or Unified Parallel C, is a parallel extension of ANSI C. UPC is developed around the distributed shared-memory programming model with constructs that can allow programmers to exploit memory locality, by placing data close to the threads that manipulate them in order to minimize remote accesses. Under the UPC memory sharing model, each thread owns a private memory and has a logical association (affinity) with a partition of the shared memory. This paper discusses an early release of UPC Bench, a benchmark designed to reveal UPC compilers performance weaknesses to uncover opportunities for compiler optimizations. The experimental results from UPC Bench over the Compaq AlphaServer SC show that UPC Bench is capable of discovering such compiler performance problems. Further, it shows that if such performance pitfalls are avoided through compiler optimizations, distributed shared memory programming paradigms can result in high-performance, while the ease of programming is enjoyed.
Tarek A. El-Ghazawi, Sébastien Chauvin
ICPP1
2001 Mapping tasks onto nodes: a parallel local neighborhood approach
S. Mounir Alaoui, Tarek A. El-Ghazawi, Ophir Frieder, Abdelghani Bellaachia, Amine Bensaid
Future Gener. Comput. Syst.2
2001 2-phase GA-based image registration on parallel clusters
Prachya Chalermwat, Tarek A. El-Ghazawi, Jacqueline LeMoigne-Stewart
Future Gener. Comput. Syst.2
2000 Parallel mining of association rules with a Hopfield type neural network
abstract
Association rule mining (ARM) is one of the data mining problems receiving a great deal of attention in the database community. The main computation step in an ARM algorithm is frequent itemset discovery. In this paper, a frequent itemset discovery algorithm based on the Hopfield model is presented.
Jaafar Gaber, Jacques M. Bahi, Tarek A. El-Ghazawi
ICTAI3
2000 Characterizing and representing workloads for parallel computer architectures
Abdullah I. AlMojel, Tarek A. El-Ghazawi, Thomas L. Sterling
J. Syst. Archit.2
1999 Multi-Resolution Image Registration Using Genetics
abstract
In remote sensing, image registration is to find the best transform between a reference and input images that may be different due to changes in position or altitude of or noise in the sensors. Image registration is one of the first steps in the analysis of remotely sensed images and requires high computational resources. The computation time is affected by two factors: search data size and search space. This paper describes an efficient image registration algorithm that uses multi-resolution wavelet decomposed images to reduce the search data size, and Genetic Algorithms to optimize the search solution space. Experimental results have shown subpixel accuracy and high efficiency over conventional methods.
Prachya Chalermwat, Tarek A. El-Ghazawi
ICIP (2)2
1999 Remote Data Access via the SIESIP Distributed Information System
abstract
Illustrates a distributed system that provides online searching, analysis and ordering capabilities for distributed Earth science data. The system is under development by a consortium led by George Mason University in a project called Seasonal-to-Interannual Earth Science Information Partners (SIESIP) as a part of a federation of information partners funded by NASA. The integrated system is composed of data, a database management system (DBMS), communication protocols, data analysis tools and a user interface. Through a Web-based Java GUI, users can search the DBMS for metadata information, conduct content-based searches, perform some initial analyses and issue an order for the selected data.
Changzhou Wang, Menas Kafatos, Xiaoyang Sean Wang, Tarek A. El-Ghazawi
SSDBM5
1997 Load-Balanced Sparse Matrix--Vector Multiplication on Parallel Computers
Sorin G. Nastea, Ophir Frieder, Tarek A. El-Ghazawi
J. Parallel Distributed Comput.3
1996 Parallel Input/Output Impact on Sparse Matrix Compression
abstract
Sparse matrices efficiently store structured information, particularly when represented in compressed formats. The advantages of using compressed formats rather than expanded representations are reduced storage space and faster computation achieved by avoiding processing the zero elements. We address the I/O bottleneck associated with the compression operation. We show that such a bottleneck can be reduced if parallel I/O techniques are used. We study several available parallel file system (PFS) access modes available on an Intel Paragon with 64 processing nodes (among whom 56 are compute nodes and 3 are I/O nodes).
Sorin G. Nastea, Tarek A. El-Ghazawi, Ophir Frieder
Data Compression Conference2
1992 A Unified Approach to Fault-Tolerant Routing
abstract
A theoretical study of the connectivity and fault tolerance of Cartesian product networks is presented. The theoretical results are used to synthesize provably correct adaptive fault-tolerant algorithms from ones written for the component networks. The theoretical foundations that relate the connectivity of a Cartesian product network, the connectivity of the component networks, and the number of faulty components are established. It is shown that the connectivity of a product network is at least the sum of the connectivities of its factor networks. Based on the constructive connectivity proof, an adaptive, generic, distributed algorithm that can perform successful point-to-point routing in product networks, in the presence of faults, is devised. A proof of correctness of the algorithm is provided.>
Tarek A. El-Ghazawi, Abdou Youssef
ICDCS1