Engin Kayraklioglu

dblp:160/2340 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
3since 2021 · last 2024
0000-0002-4966-3812ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 6 first-author · 3 since 2021
YearPublicationVenuePosition
2024 Investigating Portability in Chapel for Tree-Based Optimization on GPU-Powered Clusters
Tiago Carneiro 0001, Engin Kayraklioglu, Guillaume Helbecque, Nouredine Melab
Euro-Par (3)2
2023 Virtualizing a Post-Moore's Law Analog Mesh Processor: The Case of a Photonic PDE Accelerator
abstract
Innovative processor architectures aim to play a critical role in future sustainment of performance improvements under severe limitations imposed by the end of Moore’s Law. The Reconfigurable Optical Computer (ROC) is one such innovative, Post-Moore’s Law processor. ROC is designed to solve partial differential equations in one shot as opposed to existing solutions, which are based on costly iterative computations. This is achieved by leveraging physical properties of a mesh of optical components that behave analogously to lumped electrical components. However, virtualization is required to combat shortfalls of the accelerator hardware. Namely, (1) the infeasibility of building large photonic arrays to accommodate arbitrarily large problems and (2) underutilization brought about by mismatches in problem and accelerator mesh sizes due to future advances in manufacturing technology. In this work, we introduce an architecture and methodology for lightweight virtualization of ROC that exploits advantages borne from optical computing technology. Specifically, we apply temporal and spatial virtualization to ROC and then extend the accelerator scheduling tradespace with the introduction of spectral virtualization. Additionally, we investigate multiple resource scheduling strategies for a system-on-chip (SoC)-based PDE acceleration architecture and show that virtual configuration management offers a speedup of approximately 2×. Finally, we show that overhead from virtualization is minimal, and our experimental results show two orders of magnitude increased speed as compared to microprocessor execution while keeping errors due to virtualization under 10%.
Jeff Anderson, Engin Kayraklioglu, Hamid Reza Imani, Chen Shen 0005, Mario Miscuglio, Volker J. Sorger, Tarek A. El-Ghazawi
ACM Trans. Embed. Comput. Syst.2
2021 A Machine-Learning-Based Framework for Productive Locality Exploitation
abstract
Data locality is of extreme importance in programming distributed-memory architectures due to its implications on latency and energy consumption. Automated compiler and runtime system optimization studies have attempted to improve data locality exploitation without burdening the programmer. However, due to the difficulty of static code analysis, conservatism in compiler optimizations to avoid errors, and cost of dynamic analysis, the efficacy of automated optimizations is limited. Therefore, programmers need to spend significant effort in optimizing locality while creating applications for distributed memory parallel systems. We present a machine-learning based framework to automatically exploit locality in distributed memory applications. This framework takes application source whose time-critical blocks are marked by pragmas, and produces optimized source code that uses a regressor for efficient data movement. The regressor is trained with automatically-collected application profiles with very small input data sizes. We integrate our prototype in the Chapel language stack. In our experiments, we show that the Elastic Net model is the ideal regressor for our case and applications that utilize Elastic Net can perform very similarly to programmer-optimized versions. We also show that such regressors can be trained within few minutes on a cluster or within 30 minutes on a workstation, including data collection.
Engin Kayraklioglu, Erwan Favry, Tarek A. El-Ghazawi
IEEE Trans. Parallel Distributed Syst.1
2020 Software stack for an analog mesh computer: the case of a nanophotonic PDE accelerator
abstract
The slowing of Moore's Law is forcing the computer industry to embrace domain-specific hardware, which must be coupled with general-purpose traditional systems. This architecture is most useful when large compute power is needed. Among the most compute-intensive applications is the simulation of physical sciences. To maximize productivity in this domain, a variety accelerators have been proposed; however, the analog mesh computer has consistently been proven to require the shortest time-to-solution when targeted toward the Poisson equation. Recent advances in material science have increased the flexibility of the analog mesh computer, positioning it well for future heterogeneous computing systems. However, for the analog mesh computer to gain widespread acceptance, a software stack is required to enable seamless integration with a classical computer. Here, we introduce a software stack designed for the class of analog mesh computers that efficiently generates mesh mappings of a physical problem by enabling users to describe their problem in terms of boundary conditions and mesh parameters. Experiments on a specific implementation of analog mesh computer, the nanophotonic partial differential equation accelerator, show that this stack enables problem-to-mesh scalability expected by the scientific community.
Engin Kayraklioglu, Jeff Anderson, Hamid Reza Imani, Volker J. Sorger, Tarek A. El-Ghazawi
CF1
2019 A Machine Learning Approach for Productive Data Locality Exploitation in Parallel Computing Systems
abstract
Data locality is of extreme importance in programming distributed-memory architectures due to its implications on latency and energy consumption. Automated compiler and runtime system optimization studies have attempted to improve data locality exploitation without burdening the programmer. However, due to the difficulty of static code analysis, conservatism in compiler optimizations to avoid errors, and cost of dynamic analysis, the efficacy of automated optimizations is limited. Therefore, programmers need to spend significant effort in optimizing locality. In this work, we present an automated code optimization framework that trains neural networks using application profiles for small data sizes that exhibit similar patterns to larger cases. The application is then modified to use the neural network to improve data locality exploitation. We prototype our framework for the Chapel language and integrate with the language stack. We experimentally demonstrate that our framework can learn access patterns and create optimized executables in minutes. The resulting executables perform more than one order of magnitude faster than unoptimized code, and are comparable to manual locality optimization without burdening the programmer and hindering productivity.
Engin Kayraklioglu, Erwan Favry, Tarek A. El-Ghazawi
CCGRID1
2018 APAT: an access pattern analysis tool for distributed arrays
abstract
Distributed arrays reduce programming effort through implicit communication. However, relying solely on this abstraction causes fine-grained communication and performance overhead. A variety of optimization techniques can be used to mitigate such overheads. However, these techniques require a thorough understanding of how distributed arrays are accessed which can be very challenging in realistic use cases. We present Access Pattern Analysis Tool (APAT) for distributed arrays. APAT is a framework that can be integrated into language software stack to efficiently collect access logs and analyze them. We show that APAT can help discover optimization opportunities that can lead to up to 35% improvement.
Engin Kayraklioglu, Tarek A. El-Ghazawi
CF1
2018 LAPPS: Locality-Aware Productive Prefetching Support for PGAS
abstract
Prefetching is a well-known technique to mitigate scalability challenges in the Partitioned Global Address Space (PGAS) model. It has been studied as either an automated compiler optimization or a manual programmer optimization. Using the PGAS locality awareness, we define a hybrid tradeoff. Specifically, we introduce locality-aware productive prefetching support for PGAS. Our novel, user-driven approach strikes a balance between the ease-of-use of compiler-based automated prefetching and the high performance of the laborious manual prefetching. Our prototype implementation in Chapel shows that significant scalability and performance improvements can be achieved with minimal effort in common applications.
Engin Kayraklioglu, Michael P. Ferguson, Tarek A. El-Ghazawi
ACM Trans. Archit. Code Optim.1
2017 HPC-Oriented Toolchain for Hardware Simulators
abstract
Hardware design is an essential part of research in high performance computing. Initial efforts in hardware research consist of analyzing the design ideas in a software simulator. This allows chip designers to minimize amount of manufacturing that would be too costly and to avoid doing FPGA designs which are even more time consuming. Simulating a hardware design involves running many tests that try different configurations. Moreover, hardware simulators generally do not support multi-threaded simulation. This causes major scalability issues as simulated HPC architectures have increasing number of cores.In this paper, we present a front-end framework for hardware simulators that allows chip designers to create simulation recipes and run them in parallel. This way, a cluster can easily be used to parallelize the hardware simulations. Our framework is implemented in Python3 and have functions such as running unlimited configurations, cooperating with job managers such as Slurm and SGE and collecting and parsing results.
Olivier Serres, Engin Kayraklioglu, Tarek A. El-Ghazawi
CLUSTER2
2016 Exploiting Hierarchical Locality in Deep Parallel Architectures
abstract
Parallel computers are becoming deeply hierarchical. Locality-aware programming models allow programmers to control locality at one level through establishing affinity between data and executing activities. This, however, does not enable locality exploitation at other levels. Therefore, we must conceive an efficient abstraction of hierarchical locality and develop techniques to exploit it. Techniques applied directly by programmers, beyond the first level, burden the programmer and hinder productivity. In this article, we propose the Parallel Hierarchical Locality Abstraction Model for Execution (PHLAME). PHLAME is an execution model to abstract and exploit machine hierarchical properties through locality-aware programming and a runtime that takes into account machine characteristics, as well as a data sharing and communication profile of the underlying application. This article presents and experiments with concepts and techniques that can drive such runtime system in support of PHLAME. Our experiments show that our techniques scale up and achieve performance gains of up to 88%.
Ahmad Anbar, Olivier Serres, Engin Kayraklioglu, Abdel-Hameed A. Badawy, Tarek A. El-Ghazawi
ACM Trans. Archit. Code Optim.3
2015 Assessing Memory Access Performance of Chapel through Synthetic Benchmarks
abstract
The Partitioned Global Address Space(PGAS) programming model strikes a balance between high performance and locality awareness. As a PGAS language, Chapel relieves programmers from handling details of data movement in a distributed memory environment, by presenting a flat memory space that is logically partitioned among executing entities. Traversing such a space requires address mapping to the system virtual address space, and as such, this abstraction inevitably causes major overheads during memory accesses. In this paper, we analyzed the extent of this overhead by implementing a micro benchmark to test different types of memory accesses that can be observed in Chapel. We showed that, as the locality gets exploited speedup gains up to 35x can be achieved. This was demonstrated through hand tuning, however. More productive means should be provided to deliver such performance improvement without excessively burdening programmers. Therefore, we also discuss possibilities to increase Chapel's performance through standard libraries, compiler, runtime and/or hardware support to handle different types of memory accesses more efficiently.
Engin Kayraklioglu, Tarek A. El-Ghazawi
CCGRID1