Ami Marowka

dblp:81/5993 · DBLP profile ↗
← Back
22ranked-venue papers
22as first author
9since 2021 · last 2026
0000-0003-0914-2024ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 16 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Evaluating SYCL as a unified programming model for heterogeneous systems
Ami Marowka
Parallel Comput.1
2025 Portability efficiency approach for calculating performance portability
Ami Marowka
Future Gener. Comput. Syst.1
2025 Special issue on advances in techniques for assessment performance portability of HPC applications
abstract
This special issue aims to present new developments and advances in techniques for assessment performance portability of high performance computing applications. It contains revised and extended versions of selected papers presented at the 10th Workshop on Language-Based Parallel Programming Models, WLPP 2024, which was a part of 15th International Conference on Parallel Processing and Applied Mathematics, PPAM 2024, held on September 8–11, 2024, in Ostrava, Czech Republic.
Ami Marowka, Przemyslaw Stpiczynski, Roman Wyrzykowski
Future Gener. Comput. Syst.1
2023 Toward Open Repository of Performance Portability of Applications, Benchmarks and Models
abstract
The adoption of heterogeneous computing systems based on diverse architectures to achieve exascale computing power has worsened the performance portability problem of scientific applications that were designed to run on these platforms. To cope with the challenges posed by supercomputing, new performance portability frameworks have been developed along-side advanced methods and metrics to evaluate the performance portability of heterogeneous applications. However, many studies have shown that the new methods and metrics do not produce coherent results which yield clear conclusions that are required for designing the hardware and software architectures of tomorrow's supercomputing systems. We outline a proposal to establish an open repository of performance portability of applications, benchmarks and models which will be standardized, objective, and based on strict operating and reporting guidelines. Such guidelines will ensure a fair, comparable and meaningful measure of the performance portability while the requirement for a detailed disclosure of the obtained results and the configuration settings will ensure the reproducibility of the reported results.
Ami Marowka
SBAC-PAD1
2023 A comparison of two performance portability metrics
abstract
Summary The rise in the demand for new performance portability frameworks for heterogeneous computing systems has brought with it a number of proposals of workable metrics for evaluating the performance portability of applications. The aim of this article is twofold. First, we analyze the underlying principles of the criteria and definition of the revised Ⴔ metric and show that these principles are partially correct and suffer from a lack of a solid and clear performance portability model. We prove mathematically and demonstrate it practically that the principles are only correct for the architectural efficiency approach based on throughputs but are incorrect for the popular application efficiency approach based on run‐times. Second, we are examining whether the Ⴔ and metrics meet the requirements of consistency, proportionality, and lossless information. We use examples from the scientific literature to show the reader that the Ⴔ metric loses information while the metric and other metrics do not lose information.
Ami Marowka
Concurr. Comput. Pract. Exp.1
2023 Editorial on Advances in High Performance Programming
Ami Marowka, Przemyslaw Stpiczynski
Parallel Comput.1
2022 On the Performance Portability of OpenACC, OpenMP, Kokkos and RAJA
abstract
Performance Portability frameworks are becoming more central and essential in heterogeneous computing systems. However, the developer toolbox lacks the tools to assess the performance portability degree of these frameworks.
Ami Marowka
HPC Asia1
2022 Reformulation of the performance portability metric
abstract
Summary The 3‐P challenge of high‐performance programming—performance, portability and productivity—has become more difficult than ever in the age of heterogeneous computing. It would be naïve to think that the performance portability problem can be completely solved, but it can certainly be reduced and made tolerable. However, first and foremost, an agreement is needed on what it means for an application to be performance portable. Unfortunately, there is still no consensus in the scientific community on a workable definition of the term performance portability. Several years ago, a comprehensive effort was made to formulate a novel definition of performance portability and an associated metric. Since the new metric was first introduced, it has been widely adopted by the scientific community, and many advanced studies have used it. Unfortunately, the definition of the new metric has flaws. This article presents a proof of the theoretical flaws in the definition of the new metric, considers the practical implications of these flaws as reflected in many studies that have used it in recent years, and proposes a revised metric that addresses the flaws and provides guidelines on how to use it correctly.
Ami Marowka
Softw. Pract. Exp.1
2021 Toward a Better Performance Portability Metric
abstract
Several years ago, a comprehensive effort was made to formulate a novel definition of performance portability and an associated metric. Since the new metric was first introduced, it has been widely adopted by the scientific community, and many advanced studies have used it. Unfortunately, the definition of the new metric has flaws. This article presents a proof of the theoretical flaws in the definition of the new metric, considers the practical implications of these flaws as reflected in many studies that have used it in recent years, and proposes a revised metric that addresses the flaws.
Ami Marowka
PDP1
2020 On the performance difference between theory and practice for parallel algorithms
Ami Marowka
J. Parallel Distributed Comput.1
2020 Editorial on the special issue on advances in parallel programming: Languages, models and algorithms
Ami Marowka, Przemyslaw Stpiczynski
J. Parallel Distributed Comput.1
2018 Python accelerators for high-performance computing
Ami Marowka
J. Supercomput.1
2018 Special section on parallel programming
abstract
WLPP 2017 was a two-day workshop focusing on high-level programming for large-scale parallel systems and multicore processors, with special emphasis on component architectures and models.Its goal was to bring together researchers working in the areas of applications, computational models, language design, compilers, system architecture, and programming tools to discuss new developments in programming Clouds and parallel systems.Papers in this section cover the most important topics presented during the workshop.The first two deal with parallel programming models.Thoman et al. [4] provide an initial task-focused taxonomy for HPC technologies, which covers both programming interfaces and runtime mechanisms and discuss its usefulness by classifying state-of-the-art task-based environments that are used today.Posner and Fohry [7] propose a hybrid work stealing scheme, which combines the lifeline-based variant of distributed task pools with the node-internal load balancing implemented as an extension of the APGAS library for Java.
Ami Marowka, Przemyslaw Stpiczynski
J. Supercomput.1
2014 Energy-Efficient Management of DVFS-Enabled Integrated Microprocessors
abstract
This research presents analytical models based on an energy consumption metric to analyze the impact of dynamic frequency scaling on the energy consumption of various architectural design choices for hybrid architecture chips. The power consumption implications of different processing schemes and various chip configurations were also analyzed. The analysis shows that by choosing the optimal hardware configuration, the energy savings can be increased considerably while keeping sacrifices in performance at tolerable levels.
Ami Marowka
PDP1
2014 Maximizing energy saving of dual-architecture processors using DVFS
Ami Marowka
J. Supercomput.1
2012 Energy Consumption Modeling for Hybrid Computing
Ami Marowka
Euro-Par1
2012 Extending Amdahl's Law for Heterogeneous Computing
abstract
Energy will be a major limiting factor in future multi-core architectures, so optimizing performance per watt should be a key driver for next generation massive-core architectures. Recent studies show that heterogeneous chips integrating different core architectures, such as CPU and GPU, on a single die is the most promising solution. We investigated how energy efficiency and scalability are affected by the power constraints imposed on contemporary hybrid CPU-GPU processors. Analytical models were developed to extend Amdahl's Law by accounting for energy limitations before examining the three processing modes available to heterogeneous processors, i.e., symmetric, asymmetric, and simultaneous asymmetric. The analysis shows clearly that greater parallelism is the most important factor affecting power consumption.
Ami Marowka
ISPA1
2008 Performance of OpenMP Benchmarks on Multicore Processors
Ami Marowka
ICA3PP1
2008 BSP2OMP: A compiler for translating BSP programs to OpenMP
abstract
The convergence of the two widely used parallel programming paradigms, shared- memory and distributed- shared-memory parallel programming models, into a unified parallel programming model is crucial for parallel computing to become the next mainstream programming paradigm. We study the design differences and the performance issues of two parallel programming models: a shared- memory programming model (OpenMP) and a distributed- shared programming model (BSP). The study was carried out by designing a compiler for translating BSP parallel programs to an OpenMP programming model called BSP20MP. Analysis of the compiler outcome, and of the performance of the compiled programs, show that the two models are based on very similar underlying principles and mechanisms.
Ami Marowka
IPDPS1
2007 Routing Speedup in Multicore-Based Ad Hoc Networks
abstract
The integration of multicore processors into wireless mobile devices is creating new opportunities to enhance the speed and scalability of message routing in ad hoc networks. In this paper we study the impact of multicore technology on routing speed and node efficiency, and draw conclusions regarding the measures that should be taken to conserve energy and prolong the lifetime of a network. We formally define three metrics and use them for performance evaluation: time-to-destination (T2D), average routing speedup (ARS), and average-node-efficiency (ANE). The T2D metric is the time a message takes to travel to its destination in a loaded traffic network. ARS measures the average routing speed gained by a multicore-based network over a single-core based network, and ANE measures the average efficiency of a node, or the number of active cores. These benchmarks show that routing speedup in networks with multicore nodes increases linearly with the number of cores and significantly decrease traffic bottlenecks, while allowing more routings to be executed simultaneously. The average node efficiency, however, decreases linearly with the number of cores per node. Power-aware protocols and energy management techniques should therefore be developed to turn off the unused cores.
Ami Marowka
ISPDC1
2006 Power-dependable transactions in mobile networks
abstract
We define a quality-of-power-service (QoPS) metric to evaluate the efficiency of power-aware routing protocols in wireless ad-hoc networks. The aim of power management of routing protocols is to prolong the life-time of individual nodes in wireless network and thus to increase the delivery rate of unicast transactions. QoPS metric is applied to different location-based unicast transaction protocols. The results confirm that power-relative distribution of data streams in multi-paths unicast transaction protocols consume substantially less energy from individual nodes than from other distribution methods. The locality distribution phenomenon discovered by the simulations explains, on the one hand, the long lifetime of large, dense, and highly degree wireless networks, and on the other hand, the short lifetime of small, sparse, and low degree networks.
Ami Marowka, David Semé
IPDPS1
2004 OpenMP-oriented applications for distributed shared memory architectures
abstract
Abstract The rapid rise of OpenMP as the preferred parallel programming paradigm for small‐to‐medium scale parallelism could slow unless OpenMP can show capabilities for becoming the model‐of‐choice for large scale high‐performance parallel computing in the coming decade. The main stumbling block for the adaptation of OpenMP to distributed shared memory (DSM) machines, which are based on architectures like cc‐NUMA, stems from the lack of capabilities for data placement among processors and threads for achieving data locality. The absence of such a mechanism causes remote memory accesses and inefficient cache memory use, both of which lead to poor performance. This paper presents a simple software programming approach called copy‐inside–copy‐back (CC) that exploits the data privatization mechanism of OpenMP for data placement and replacement. This technique enables one to distribute data manually without taking away control and flexibility from the programmer and is thus an alternative to the automat and implicit approaches. Moreover, the CC approach improves on the OpenMP‐SPMD style of programming that makes the development process of an OpenMP application more structured and simpler. The CC technique was tested and analyzed using the NAS Parallel Benchmarks on SGI Origin 2000 multiprocessor machines. This study shows that OpenMP improves performance of coarse‐grained parallelism, although a fast copy mechanism is essential. Copyright © 2004 John Wiley & Sons, Ltd.
Ami Marowka, Zhenying Liu, Barbara M. Chapman
Concurr. Comput. Pract. Exp.1