EDBT 2026 Demo / reviewers in the wild / expert
Ami Marowka
dblp:81/5993
· DBLP profile ↗
22ranked-venue papers
22as first author
9since 2021 · last 2026
0000-0003-0914-2024ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 16 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating SYCL as a unified programming model for heterogeneous systems
Ami Marowka |
Parallel Comput. | 1 |
| 2025 | Portability efficiency approach for calculating performance portability
Ami Marowka |
Future Gener. Comput. Syst. | 1 |
| 2025 | Special issue on advances in techniques for assessment performance portability of HPC applicationsabstractThis special issue aims to present new developments and advances in techniques for assessment performance portability of high performance computing applications. It contains revised and extended versions of selected papers presented at the 10th Workshop on Language-Based Parallel Programming Models, WLPP 2024, which was a part of 15th International Conference on Parallel Processing and Applied Mathematics, PPAM 2024, held on September 8–11, 2024, in Ostrava, Czech Republic. Ami Marowka, Przemyslaw Stpiczynski, Roman Wyrzykowski |
Future Gener. Comput. Syst. | 1 |
| 2023 | Toward Open Repository of Performance Portability of Applications, Benchmarks and ModelsabstractThe adoption of heterogeneous computing systems based on diverse architectures to achieve exascale computing power has worsened the performance portability problem of scientific applications that were designed to run on these platforms. To cope with the challenges posed by supercomputing, new performance portability frameworks have been developed along-side advanced methods and metrics to evaluate the performance portability of heterogeneous applications. However, many studies have shown that the new methods and metrics do not produce coherent results which yield clear conclusions that are required for designing the hardware and software architectures of tomorrow's supercomputing systems. We outline a proposal to establish an open repository of performance portability of applications, benchmarks and models which will be standardized, objective, and based on strict operating and reporting guidelines. Such guidelines will ensure a fair, comparable and meaningful measure of the performance portability while the requirement for a detailed disclosure of the obtained results and the configuration settings will ensure the reproducibility of the reported results. Ami Marowka |
SBAC-PAD | 1 |
| 2023 | A comparison of two performance portability metricsabstractSummary The rise in the demand for new performance portability frameworks for heterogeneous computing systems has brought with it a number of proposals of workable metrics for evaluating the performance portability of applications. The aim of this article is twofold. First, we analyze the underlying principles of the criteria and definition of the revised Ⴔ metric and show that these principles are partially correct and suffer from a lack of a solid and clear performance portability model. We prove mathematically and demonstrate it practically that the principles are only correct for the architectural efficiency approach based on throughputs but are incorrect for the popular application efficiency approach based on run‐times. Second, we are examining whether the Ⴔ and metrics meet the requirements of consistency, proportionality, and lossless information. We use examples from the scientific literature to show the reader that the Ⴔ metric loses information while the metric and other metrics do not lose information. Ami Marowka |
Concurr. Comput. Pract. Exp. | 1 |
| 2023 | Editorial on Advances in High Performance Programming
Ami Marowka, Przemyslaw Stpiczynski |
Parallel Comput. | 1 |
| 2022 | On the Performance Portability of OpenACC, OpenMP, Kokkos and RAJAabstractPerformance Portability frameworks are becoming more central and essential in heterogeneous computing systems. However, the developer toolbox lacks the tools to assess the performance portability degree of these frameworks. Ami Marowka |
HPC Asia | 1 |
| 2022 | Reformulation of the performance portability metricabstractSummary The 3‐P challenge of high‐performance programming—performance, portability and productivity—has become more difficult than ever in the age of heterogeneous computing. It would be naïve to think that the performance portability problem can be completely solved, but it can certainly be reduced and made tolerable. However, first and foremost, an agreement is needed on what it means for an application to be performance portable. Unfortunately, there is still no consensus in the scientific community on a workable definition of the term performance portability. Several years ago, a comprehensive effort was made to formulate a novel definition of performance portability and an associated metric. Since the new metric was first introduced, it has been widely adopted by the scientific community, and many advanced studies have used it. Unfortunately, the definition of the new metric has flaws. This article presents a proof of the theoretical flaws in the definition of the new metric, considers the practical implications of these flaws as reflected in many studies that have used it in recent years, and proposes a revised metric that addresses the flaws and provides guidelines on how to use it correctly. Ami Marowka |
Softw. Pract. Exp. | 1 |
| 2021 | Toward a Better Performance Portability MetricabstractSeveral years ago, a comprehensive effort was made to formulate a novel definition of performance portability and an associated metric. Since the new metric was first introduced, it has been widely adopted by the scientific community, and many advanced studies have used it. Unfortunately, the definition of the new metric has flaws. This article presents a proof of the theoretical flaws in the definition of the new metric, considers the practical implications of these flaws as reflected in many studies that have used it in recent years, and proposes a revised metric that addresses the flaws. Ami Marowka |
PDP | 1 |
| 2020 | On the performance difference between theory and practice for parallel algorithms
Ami Marowka |
J. Parallel Distributed Comput. | 1 |
| 2020 | Editorial on the special issue on advances in parallel programming: Languages, models and algorithms
Ami Marowka, Przemyslaw Stpiczynski |
J. Parallel Distributed Comput. | 1 |
| 2018 | Python accelerators for high-performance computing
Ami Marowka |
J. Supercomput. | 1 |
| 2018 | Special section on parallel programmingabstractWLPP 2017 was a two-day workshop focusing on high-level programming for large-scale parallel systems and multicore processors, with special emphasis on component architectures and models.Its goal was to bring together researchers working in the areas of applications, computational models, language design, compilers, system architecture, and programming tools to discuss new developments in programming Clouds and parallel systems.Papers in this section cover the most important topics presented during the workshop.The first two deal with parallel programming models.Thoman et al. [4] provide an initial task-focused taxonomy for HPC technologies, which covers both programming interfaces and runtime mechanisms and discuss its usefulness by classifying state-of-the-art task-based environments that are used today.Posner and Fohry [7] propose a hybrid work stealing scheme, which combines the lifeline-based variant of distributed task pools with the node-internal load balancing implemented as an extension of the APGAS library for Java. Ami Marowka, Przemyslaw Stpiczynski |
J. Supercomput. | 1 |
| 2014 | Energy-Efficient Management of DVFS-Enabled Integrated MicroprocessorsabstractThis research presents analytical models based on an energy consumption metric to analyze the impact of dynamic frequency scaling on the energy consumption of various architectural design choices for hybrid architecture chips. The power consumption implications of different processing schemes and various chip configurations were also analyzed. The analysis shows that by choosing the optimal hardware configuration, the energy savings can be increased considerably while keeping sacrifices in performance at tolerable levels. Ami Marowka |
PDP | 1 |
| 2014 | Maximizing energy saving of dual-architecture processors using DVFS
Ami Marowka |
J. Supercomput. | 1 |
| 2012 | Energy Consumption Modeling for Hybrid Computing
Ami Marowka |
Euro-Par | 1 |
| 2012 | Extending Amdahl's Law for Heterogeneous ComputingabstractEnergy will be a major limiting factor in future multi-core architectures, so optimizing performance per watt should be a key driver for next generation massive-core architectures. Recent studies show that heterogeneous chips integrating different core architectures, such as CPU and GPU, on a single die is the most promising solution. We investigated how energy efficiency and scalability are affected by the power constraints imposed on contemporary hybrid CPU-GPU processors. Analytical models were developed to extend Amdahl's Law by accounting for energy limitations before examining the three processing modes available to heterogeneous processors, i.e., symmetric, asymmetric, and simultaneous asymmetric. The analysis shows clearly that greater parallelism is the most important factor affecting power consumption. Ami Marowka |
ISPA | 1 |
| 2008 | Performance of OpenMP Benchmarks on Multicore Processors
Ami Marowka |
ICA3PP | 1 |
| 2008 | BSP2OMP: A compiler for translating BSP programs to OpenMPabstractThe convergence of the two widely used parallel programming paradigms, shared- memory and distributed- shared-memory parallel programming models, into a unified parallel programming model is crucial for parallel computing to become the next mainstream programming paradigm. We study the design differences and the performance issues of two parallel programming models: a shared- memory programming model (OpenMP) and a distributed- shared programming model (BSP). The study was carried out by designing a compiler for translating BSP parallel programs to an OpenMP programming model called BSP20MP. Analysis of the compiler outcome, and of the performance of the compiled programs, show that the two models are based on very similar underlying principles and mechanisms. Ami Marowka |
IPDPS | 1 |
| 2007 | Routing Speedup in Multicore-Based Ad Hoc NetworksabstractThe integration of multicore processors into wireless mobile devices is creating new opportunities to enhance the speed and scalability of message routing in ad hoc networks. In this paper we study the impact of multicore technology on routing speed and node efficiency, and draw conclusions regarding the measures that should be taken to conserve energy and prolong the lifetime of a network. We formally define three metrics and use them for performance evaluation: time-to-destination (T2D), average routing speedup (ARS), and average-node-efficiency (ANE). The T2D metric is the time a message takes to travel to its destination in a loaded traffic network. ARS measures the average routing speed gained by a multicore-based network over a single-core based network, and ANE measures the average efficiency of a node, or the number of active cores. These benchmarks show that routing speedup in networks with multicore nodes increases linearly with the number of cores and significantly decrease traffic bottlenecks, while allowing more routings to be executed simultaneously. The average node efficiency, however, decreases linearly with the number of cores per node. Power-aware protocols and energy management techniques should therefore be developed to turn off the unused cores. Ami Marowka |
ISPDC | 1 |
| 2006 | Power-dependable transactions in mobile networksabstractWe define a quality-of-power-service (QoPS) metric to evaluate the efficiency of power-aware routing protocols in wireless ad-hoc networks. The aim of power management of routing protocols is to prolong the life-time of individual nodes in wireless network and thus to increase the delivery rate of unicast transactions. QoPS metric is applied to different location-based unicast transaction protocols. The results confirm that power-relative distribution of data streams in multi-paths unicast transaction protocols consume substantially less energy from individual nodes than from other distribution methods. The locality distribution phenomenon discovered by the simulations explains, on the one hand, the long lifetime of large, dense, and highly degree wireless networks, and on the other hand, the short lifetime of small, sparse, and low degree networks. Ami Marowka, David Semé |
IPDPS | 1 |
| 2004 | OpenMP-oriented applications for distributed shared memory architecturesabstractAbstract The rapid rise of OpenMP as the preferred parallel programming paradigm for small‐to‐medium scale parallelism could slow unless OpenMP can show capabilities for becoming the model‐of‐choice for large scale high‐performance parallel computing in the coming decade. The main stumbling block for the adaptation of OpenMP to distributed shared memory (DSM) machines, which are based on architectures like cc‐NUMA, stems from the lack of capabilities for data placement among processors and threads for achieving data locality. The absence of such a mechanism causes remote memory accesses and inefficient cache memory use, both of which lead to poor performance. This paper presents a simple software programming approach called copy‐inside–copy‐back (CC) that exploits the data privatization mechanism of OpenMP for data placement and replacement. This technique enables one to distribute data manually without taking away control and flexibility from the programmer and is thus an alternative to the automat and implicit approaches. Moreover, the CC approach improves on the OpenMP‐SPMD style of programming that makes the development process of an OpenMP application more structured and simpler. The CC technique was tested and analyzed using the NAS Parallel Benchmarks on SGI Origin 2000 multiprocessor machines. This study shows that OpenMP improves performance of coarse‐grained parallelism, although a fast copy mechanism is essential. Copyright © 2004 John Wiley & Sons, Ltd. Ami Marowka, Zhenying Liu, Barbara M. Chapman |
Concurr. Comput. Pract. Exp. | 1 |