VLDB 2026 Research / reviewers in the wild / expert
Gihan R. Mudalige
dblp:26/1059 · also Gihan Ravideva Mudalige
· DBLP profile ↗
28ranked-venue papers
6as first author
11since 2021 · last 2025
0000-0002-1398-5174ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 6 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Automatic Code-Generation for Accelerating Structured-Mesh-Based Explicit Numerical Solvers on FPGAsabstractStructured-mesh-based stencil computations are a common motif in many numerical algorithms, such as for solving PDEs. Recent work has shown promising runtime performance and energy efficiency when mapping these applications to FPGAs. However, this requires significant manual effort with hardwarespecific optimizations and customization. We present a new codegeneration framework that automates the mapping of stencil applications to FPGAs. From a domain-specific declaration, the framework applies radical, ad-hoc, and dynamic optimizations, including a novel window-buffer chaining scheme, and an on-chip loopback pipeline approach that minimizes memory transaction overhead. We show that the generated code is able to match or exceed the performance of hand-tuned state-of-the-art designs. A range of stencil solvers benchmarked, including non-trivial multi-stencil solvers, on two FPGAs, demonstrating performanceportability. Comparisons to best-in-class solvers on an Nvidia H100 GPU indicate competitive performance with 2-19× energy savings. Beniel Thileepan, Suhaib A. Fahmy, Gihan R. Mudalige |
PACT | 3 |
| 2025 | The P3 Explorer: An Open Database of Performance, Portability, and ProductivityabstractThis paper introduces a web-based tool designed to organise and present visual representations of performance, portability, and productivity (P3) data from previously published scientific studies. The P3 Explorer operates as both an open repository of application performance data and a data dashboard, providing visual heuristic analyses of performance portability and developer productivity, created using Intel’s P3 Analysis library. The goal of the project is to create a community-led database of P3 studies to better inform application developers of alternative approaches to developing new applications that target high performance on diverse hardware, considering the productivity of the developer. In this paper, we evaluate our tool using a recently published study outlining a performance portable domain-specific language for particle-in-cell applications. Steven A. Wright 0001, Zaman Lantra, Gihan R. Mudalige |
PDP | 4 |
| 2024 | Mini-Combust - An Open-Source Unstructured FGM Combustion Mini-App for Co-Designing Aero-Engines at Extreme ScaleabstractWe present Mini-Combust, a representative mini-application to explore algorithms of interest for simulating combustion in aero turbine engines. Mini-Combust uses an asynchronous coupled Lagrangian-Eulerian (particle and FVM flow field solver) approach and incorporates basic models for heat transfer, turbulence, boundary conditions, convection schemes, fuel, spray atomisation and emissions. The mini-app is developed in a modular and scalable setup with highly optimized code-paths for executing the solvers on multi-core CPU systems and on CPU/GPU heterogeneous systems. We investigate its performance and scalability on two high-performance computing systems, exploring key performance profiles on up to 32k CPU cores and 128 GPUs, solving a test case representative of a bluff-body swirl burner from industry. Results demonstrate the highly memory-bound nature of the key kernels of the solvers. We see a speedup of 8.8× on an H100 node compared to a power-equivalent number of CPU nodes when offloading the most time-consuming component, the Eulerian solver, to GPUs, executing the asynchronous solvers in a hybrid CPU-GPU manner on an NVIDIA H100 GPU node. Additionally, we see how the communication overhead between the two solvers becomes more pronounced at increasing scale. Mini-Combust is available as open-source software. Samuel Curtis, Harry Waugh, Tom Deakin, Gihan R. Mudalige |
HiPC | 4 |
| 2024 | OP-PIC - an Unstructured-Mesh Particle-in-Cell DSL for Developing Nuclear Fusion SimulationsabstractParticle-in-Cell (PIC) applications form a core simulation component for designing fusion reactors and their efficient use. In this work, we introduce OP-PIC, a new embedded Domain Specific Language (DSL) for developing unstructured-mesh PIC applications. The DSL is aimed at gaining performance portability for PIC codes on current and emerging, massively parallel architectures. We investigate and bring together the state-of-the-art in PIC solver parallelization techniques, refactoring them within a multi-layered DSL. OP-PIC use source-to-source translation to generate platform-specific optimizations. These parallelizations can be reused for any application declared using the DSL’s high-level API. We showcase the performance and portability of two non-trivial PIC applications developed with OP-PIC on multiple CPU and GPU clusters, employing a number of parallelization techniques, including OpenMP, CUDA, HIP and their combinations with distributed memory (MPI) parallelization. We benchmark the OP-PIC generated code on a range of single node systems and a number of distributed-memory systems, including an AMD EPYC CPU-based HPE-Cray EX cluster, an NVIDIA V100 GPU cluster, and an AMD MI250X GPU cluster, exploring both single node and scaling performance. Results demonstrate the flexibility of the DSL to implement radically different optimizations for each platform, showing between 1.4 × to 3.5 × speed-ups with GPUs compared to CPUs on power equivalent systems and good weak-scaling to over 10 billion particles. Zaman Lantra, Steven A. Wright 0001, Gihan R. Mudalige |
ICPP | 3 |
| 2024 | Performance-Portable Multiphase Flow Solutions with Discontinuous Galerkin MethodsabstractWe present a performance portable solver workflow for developing multiphase flow simulations based on the OP2 domain-specific language. The workflow brings together the state-of-the-art in multiphase flow solutions and uses discontinuous Galerkin (DG) methods for achieving high-order accuracy within OP2’s established performance portable framework. A key aspect of the modelling is managing the fluid interface through high-order methods while maintaining stability, accuracy and performance. The new OP2-DG methods are used for developing the non-trivial multiphase liquid whistle problem, 4M to 128M element meshes / 80M to 2.4B degrees of freedom, representative of an application used in formulated products manufacturing. The OP2-generated code is benchmarked on an AMD EPYC CPU-based HPE-Cray EX cluster, a NVIDIA V100 GPU cluster and an AMD MI250X GPU cluster. The runtime and scaling performance is analysed, demonstrating how superior high-order solutions can be achieved without sacrificing performance portability while approaching tractable times-to-solution for large multiphase flow simulations that were previously out of reach. Tobias S. Flynn, Robert Sawko, Gihan R. Mudalige |
IPDPS | 3 |
| 2023 | Communication-Avoiding Optimizations for Large-Scale Unstructured-Mesh Applications with OP2abstractIn this paper, we investigate data movement-reducing and communication-avoiding optimizations and their practicable implementation for large-scale unstructured-mesh applications. Utilizing the high-level abstraction of the OP2 DSL for the unstructured-mesh class of codes, we reason about techniques for reduced communications across a consecutive sequence of loops – a loop-chain. The careful trade-off with increased redundant computation in place of data movement is analyzed for distributed-memory parallelization. A new communication-avoiding (CA) back-end for OP2 is designed, codifying these techniques such that they can be applied automatically to any OP2 application. The back-end is extended to operate on a cluster of GPUs, integrating GPU-to-GPU communication with CUDA, in combination with MPI. The new CA back-end is applied automatically to two non-trivial applications, including the OP2 version of Rolls-Royce’s production CFD application, Hydra. Performance is investigated on both CPU and GPU clusters on representative problems of 8M and 24M node mesh sizes. Results demonstrate how for select configurations the new CA back-end provides between 30 – 65% runtime reductions for the loop-chains in these applications for the mesh sizes on both an HPE Cray EX system and an NVIDIA V100 GPU cluster. We model and examine the determinants and characteristics of a given unstructured-mesh loop-chain that can lead to performance benefits with CA techniques, providing insights into the general feasibility and profitability of using the optimizations for this class of applications. Suneth Dasantha Ekanayake, István Z. Reguly, Fabio Luporini, Gihan R. Mudalige |
ICPP | 4 |
| 2023 | Predictive Analysis of Code Optimisations on Large-Scale Coupled CFD-Combustion Simulations using the CPX Mini-AppabstractAs the complexity of multi-physics simulations increases, there is a need for efficient flow of information between components. Discrete ‘coupler’ codes can abstract away this process, improving solver interoperability. One such multi-physics problem is modelling a gas turbine aero engine, where instances of rotor/stator CFD and combustion simulations are coupled. Allocating resources correctly and efficiently during production simulations is a significant challenge due to the large HPC resources required and the varying scalability of specific components, a result of differences between solver physics. In this research, we develop a coupled mini-app simulation and an accompanying performance model to help support this process. We integrate an existing Particle-In-Cell mini-app, SIMPIC, as a ‘performance proxy’ for production combustion codes in industry, into a coupled mini-app CFD simulation using the CPX mini-coupler. The bottlenecks of the workload are examined, and the performance behavior are replicated using the mini-app. A selection of optimizations are examined, allowing us to estimate the workload’s theoretical performance. The coupling of mini-apps is supported by an empirical performance model which is then used to load balance and predict the speedup of a full-scale compressor-combustor-turbine simulation of 1.2Bn cells, a production representative problem size. The model is validated on 40K-cores of an HPE-Cray EX system, predicting the runtime of the mini-app work-flow with over 75% accuracy. The developed coupled mini-apps and empirical model combination demonstrates how rapid design space and run-time setup exploration studies can be carried out to obtain the best performance from full-scale Combustion-CFD coupled simulations. Archie Powell, Gihan R. Mudalige |
IPDPS | 2 |
| 2022 | Towards Virtual Certification of Gas Turbine Engines With Performance-Portable SimulationsabstractWe present the large-scale, computational fluid dy-namics (CFD) simulation of a full gas-turbine engine compressor, demonstrating capability towards overcoming current limitations for virtual certification of aero-engine design. The simulation is carried out through a performance portable code-base on multi-core/many-core HPC clusters with a CFD-to-CFD coupled execution, combining an industrial CFD solver linked using custom coupler software. The application innovates in its design for performance portability through the OP2 domain specific library for the CFD components, allowing the automatic generation of highly optimized platform-specific parallelizations for both multi-core (CPU) and many-core (GPU) clusters from a single high-level source. The code is used for the simulation of a 4.58B node, full-annulus 10-row production-grade test compressor (DLR's Rig250), using a coupled sliding-plane setup on the ARCHER2 and Cirrus supercomputers at EPCC. The OP2 generated multiple parallelizations, together with optimized coupler configurations on heterogeneous/hybrid settings achieve, for the first time, execution of 1 revolution in less than 6 hours on 512 nodes of ARCHER2 (65k cores), with a parallel scaling efficiency of over 80 % compared to a 107 node run. Results indicate a speed up of the CFD suite by an order of a magnitude (≈30 x) relative to current production capability. Benchmarking and performance modelling project a time-to-solution of less than 5 hours on a cluster of 488xNVIDIA V100 GPUs, about 3x-4 x speedup over CPU clusters. The work demonstrates a step-change towards achieving virtual certification of aircraft engines with the requisite fidelity and tractable time-to-solution that was previously out of reach under production settings. Gihan R. Mudalige, István Z. Reguly, Arun Prabhakar, Dario Amirante, Leigh Lapworth, Stephen A. Jarvis |
CLUSTER | 1 |
| 2022 | High throughput multidimensional tridiagonal system solvers on FPGAsabstractWe present a high performance tridiagonal solver library for Xilinx FPGAs optimized for multiple multi-dimensional systems common in real-world applications. An analytical performance model is developed and used to explore the design space and obtain rapid performance estimates that are over 85% accurate. This library achieves an order of magnitude better performance when solving large batches of systems than previous FPGA work. A detailed comparison with a current state-of-the-art GPU library for multi-dimensional tridiagonal systems on an Nvidia V100 GPU shows the FPGA achieving competitive or better runtime and significant energy savings of over 30%. Through this design, we learn lessons about the types of applications where FPGAs can challenge the current dominance of GPUs. Kamalakkannan Kamalavasan, Gihan R. Mudalige, István Z. Reguly, Suhaib A. Fahmy |
ICS | 2 |
| 2021 | Predictive Analysis of Large-Scale Coupled CFD Simulations with the CPX Mini-AppabstractAs the complexity of multi-physics simulations increases, there is a need for efficient flow of information between components. Discrete ‘coupler’ codes can abstract away this process, improving solver interoperability. One such multi-physics problem is modelling the high pressure compressor of turbofan engines, where instances of rotor/stator CFD simulations are coupled. Configuring couplers and allocating resources correctly can be challenging for such problems due to the sliding interfaces between codes. In this research, we present CPX, a mini-coupler designed to model the performance behaviour of a production coupler framework at Rolls-Royce plc., used for coupling rotor/stator simulations. CPX, the first mini-coupler framework of its kind, is combined with a CFD mini-app to predict the run-time and scaling behaviour of large scale coupled CFD simulations. We demonstrate high qualitative and quantitative predictive accuracy with a less than 17 % mean error. A performance model is developed to predict the ‘optimum’ configuration of resources, and is tested to show the high accuracy of these predictions. The model is also used to project the ‘optimum’ configuration for a 6 Billion cell test case, a problem size representative of current leading-edge production workloads, on a 100,000 core cluster and a 400 GPU cluster. Further testing reveals that the ‘optimum’ configuration is unstable if not set up correctly, and therefore a trade-off needs to be made with a marginally less-than-optimal setup to ensure stability. The work illustrates the significant utility of CPX to carry out such rapid design space and run-time setup exploration studies to obtain the best performance from production CFD coupled simulations. Archie Powell, K. Choudry, Arun Prabhakar, István Z. Reguly, Dario Amirante, Stephen A. Jarvis, Gihan R. Mudalige |
HiPC | 7 |
| 2021 | High-Level FPGA Accelerator Design for Structured-Mesh-Based Explicit Numerical SolversabstractThis paper presents a workflow for synthesizing near-optimal FPGA implementations of structured-mesh based stencil applications for explicit solvers. It leverages key characteristics of the application class and its computation-communication pattern and the architectural capabilities of the FPGA to accelerate solvers for high-performance computing applications. Key new features of the workflow are (1) the unification of standard state-of-the-art techniques with a number of high-gain optimizations such as batching and spatial blocking/tiling, motivated by increasing throughput for real-world workloads and (2) the development and use of a predictive analytical model to explore the design space, and obtain resource and performance estimates. Three representative applications are implemented using the design workflow on a Xilinx Alveo U280 FPGA, demonstrating near-optimal performance and over 85% predictive model accuracy. These are compared with equivalent highly-optimized implementations of the same applications on modern HPC-grade GPUs (Nvidia V100), analyzing time to solution, bandwidth, and energy consumption. Performance results indicate comparable runtimes with the V100 GPU, with over 2× energy savings for the largest non-trivial application on the FPGA. Our investigation shows the challenges of achieving high performance on current generation FPGAs compared to traditional architectures. We discuss determinants for a given stencil code to be amenable to FPGA implementation, providing insights into the feasibility and profitability of a design and its resulting performance. Kamalakkannan Kamalavasan, Gihan R. Mudalige, István Z. Reguly, Suhaib A. Fahmy |
IPDPS | 2 |
| 2020 | Bitwise Reproducible task execution on unstructured mesh applicationsabstractMany mesh applications use floating point arithmetic which do not necessarily hold the associative laws of algebra. This could cause the application to become unreproducible. In this paper we present some work on generating a method for unstructured mesh applications to provide bitwise reproducibility between separate runs, even if they are started with different number of MPI processes. We implement our work in the OP2 domain-specific library, which provides an API that abstracts the solution of unstructured mesh computations. We carry out a performance analysis of our method applied on two applications: a simple airfoil application, and a more complex Aero application which uses a finite element method and a conjugate-gradient algorithm. We show a 2.37×to 1.49× slowdown on this applications as a price for full bitwise reproducibility. Bálint Siklósi, István Z. Reguly, Gihan R. Mudalige |
CCGRID | 3 |
| 2019 | Large-scale performance of a DSL-based multi-block structured-mesh application for Direct Numerical Simulation
Gihan R. Mudalige, István Z. Reguly, Satya P. Jammy, Christian T. Jacobs, Michael B. Giles, Neil D. Sandham |
J. Parallel Distributed Comput. | 1 |
| 2019 | Improving resilience of scientific software through a domain-specific approach
István Z. Reguly, Gihan R. Mudalige, Michael B. Giles, S. Maheswaran 0003 |
J. Parallel Distributed Comput. | 2 |
| 2019 | Locality optimized unstructured mesh algorithms on GPUs
András Attila Sulyok, Gábor Dániel Balogh, István Z. Reguly, Gihan R. Mudalige |
J. Parallel Distributed Comput. | 4 |
| 2018 | Loop Tiling in Large-Scale Stencil Codes at Run-Time with OPSabstractThe key common bottleneck in most stencil codes is data movement, and prior research has shown that improving data locality through optimisations that optimise across loops do particularly well. However, in many large PDE applications it is not possible to apply such optimisations through compilers because there are many options, execution paths and data per grid point, many dependent on run-time parameters, and the code is distributed across different compilation units. In this paper, we adapt the data locality improving optimisation called tiling for use in large OPS applications both in shared-memory and distributed-memory systems, relying on run-time analysis and delayed execution. We evaluate our approach on a number of applications, observing speedups of 2× on the Cloverleaf 2D/3D proxy applications, which contain 83(2D)/141(3D) loops, 3.5× on the linear solver TeaLeaf, and 1.7× on the compressible Navier-Stokes solver OpenSBLI. We demonstrate strong and weak scalability on up to 4608 cores of CINECA's Marconi supercomputer. We also evaluate our algorithms on Intel's Knights Landing, demonstrating maintained throughput as the problem size grows beyond 16GB, and we do scaling studies up to 8704 cores. The approach is generally applicable to any stencil DSL that provides per loop nest data access information. István Z. Reguly, Gihan R. Mudalige, Michael B. Giles |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2017 | Achieving Performance Portability for a Heat Conduction Solver Mini-Application on Modern Multi-core SystemsabstractModernizing production-grade, often legacy applications to take advantage of modern multi-core and many-core architectures can be a difficult and costly undertaking. This is especially true currently, as it is unclear which architectures will dominate future systems. The complexity of these codes can mean that parallelisation for a given architecture requires significant re-engineering. One way to assess the benefit of such an exercise would be to use mini-applications that are representative of the legacy programs.In this paper, we investigate different implementations of TeaLeaf, a mini-application from the Mantevo suite that solves the linear heat conduction equation. TeaLeaf has been ported to use many parallel programming models, including OpenMP, CUDA and MPI among others. It has also been re-engineered to use the OPS embedded DSL and template libraries Kokkos and RAJA. We use these different implementations to assess the performance portability of each technique on modern multi-core systems.While manually parallelising the application targeting and optimizing for each platform gives the best performance, this has the obvious disadvantage that it requires the creation of different versions for each and every platform of interest. Frameworks such as OPS, Kokkos and RAJA can produce executables of the program automatically that achieve comparable portability. Based on a recently developed performance portability metric, our results show that OPS and RAJA achieve an application performance portability score of 71% and 77% respectively for this application. Richard O. Kirk, Gihan R. Mudalige, István Z. Reguly, Steven A. Wright 0001, Matt Martineau, Stephen A. Jarvis |
CLUSTER | 2 |
| 2016 | Vectorizing unstructured mesh computations for many-core architecturesabstractSummary Achieving optimal performance on the latest multi‐core and many‐core architectures increasingly depends on making efficient use of the hardware's vector units. This paper presents results on achieving high performance through vectorization on CPUs and the Xeon‐Phi on a key class of irregular applications: unstructured mesh computations. Using single instruction multiple thread (SIMT) and single instruction multiple data (SIMD) programming models, we show how unstructured mesh computations map to OpenCL or vector intrinsics through the use of code generation techniques in the OP2 Domain Specific Library and explore how irregular memory accesses and race conditions can be organized on different hardware. We benchmark Intel Xeon CPUs and the Xeon‐Phi, using a tsunami simulation and a representative CFD benchmark. Results are compared with previous work on CPUs and NVIDIA GPUs to provide a comparison of achievable performance on current many‐core systems. We show that auto‐vectorization and the OpenCL SIMT model do not map efficiently to CPU vector units because of vectorization issues and threading overheads. In contrast, using SIMD vector intrinsics imposes some restrictions and requires more involved programming techniques but results in efficient code and near‐optimal performance, two times faster than non‐vectorized code. We observe that the Xeon‐Phi does not provide good performance for these applications but is still comparable with a pair of mid‐range Xeon chips. Copyright © 2015 John Wiley & Sons, Ltd. István Z. Reguly, Endre László, Gihan R. Mudalige, Michael B. Giles |
Concurr. Comput. Pract. Exp. | 3 |
| 2016 | Acceleration of a Full-Scale Industrial CFD Application with OP2abstractHydra is a full-scale industrial CFD application used for the design of turbomachinery at Rolls Royce plc., capable of performing complex simulations over highly detailed unstructured mesh geometries. Hydra presents major challenges in data organization and movement that need to be overcome for continued high performance on emerging platforms. We present research in achieving this goal through the OP2 domain-specific high-level framework, demonstrating the viability of such a high-level programming approach. OP2 targets the domain of unstructured mesh problems and enables execution on a range of back-end hardware platforms. We chart the conversion of Hydra to OP2, and map out the key difficulties encountered in the process. Specifically we show how different parallel implementations can be achieved with an active library framework, even for a highly complicated industrial application and how different optimizations targeting contrasting parallel architectures can be applied to the whole application, seamlessly, reducing developer effort and increasing code longevity. Performance results demonstrate that not only the same runtime performance as that of the hand-tuned original code could be achieved, but it can be significantly improved on conventional processor systems, and many-core systems. Our results provide evidence of how high-level frameworks such as OP2 enable portability across a wide range of contrasting platforms and their significant utility in achieving high performance without the intervention of the application programmer. István Z. Reguly, Gihan R. Mudalige, Carlo Bertolli, Michael B. Giles, Adam Betts, Paul H. J. Kelly, David Radford |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2015 | Design and Development of Domain Specific Active Libraries with Proxy ApplicationsabstractRepresentative applications are versatile tools to evaluate new programming approaches, techniques and optimisations as a way to ensure continued high performance on future computing architectures. They make experimentation much easier before adopting changes/insights into the large scientific codes. In this paper we demonstrate the important role played by representative/proxy applications in designing and developing two high-level programming approaches: namely the OP2 and OPS domain specific (active) libraries. OP2 and OPS utilizes code generation techniques to produce automatic parallelisations from a high-level abstract problem declaration. The strategy delivers significant developer productivity to the domain scientist, while at the same time allowing computational experts to adopt the latest programming models and hardware-specific optimisations into the library and code generation tools to achieve near optimal performance. We show how representative applications have been a cornerstone in the development of OP2 and OPS and chart our experiences. In particular, we demonstrate how the range of hand-tuned optimized parallelisations of the CloverLeaf hydrodynamics mini-app allowed us to gain clear evidence that the OPS based code generated parallelisations were indeed as optimal as the hand-tuned versions. Additionally, with the use of a representative application from the CFD domain we demonstrate how the optimisations discovered and applied to proxy apps are indeed directly transferable to a large-scale industrial application at Rolls Royce plc. These results provide significant evidence into the utility of representative applications to improve productivity, enable performance portability and ultimately future-proof scientific applications. István Z. Reguly, Gihan R. Mudalige, Michael B. Giles |
CLUSTER | 2 |
| 2013 | Designing OP2 for GPU architectures
Michael B. Giles, Gihan R. Mudalige, B. Spencer, Carlo Bertolli, István Z. Reguly |
J. Parallel Distributed Comput. | 2 |
| 2013 | Design and initial performance of a high-level unstructured mesh framework on heterogeneous parallel systems
Gihan R. Mudalige, Michael B. Giles, Jeyan Thiyagalingam, István Z. Reguly, Carlo Bertolli, Paul H. J. Kelly, Anne E. Trefethen |
Parallel Comput. | 1 |
| 2012 | Performance Analysis and Optimization of the OP2 Framework on Many-Core ArchitecturesabstractThis paper presents a benchmarking, performance analysis and optimization study of the OP2 ‘active’ library, which provides an abstraction framework for the parallel execution of unstructured mesh applications. OP2 aims to decouple the scientific specification of the application from its parallel implementation, and thereby achieve code longevity and near-optimal performance through re-targeting the application to execute on different multi-core/many-core hardware. Runtime performance results are presented for a representative unstructured mesh application on a variety of many-core processor systems, including traditional X86 architectures from Intel (Xeon based on the older Penryn and current Nehalem micro-architectures) and GPU offerings from NVIDIA (GTX260, Tesla C2050). Our analysis demonstrates the contrasting performance between the use of CPU (OpenMP) and GPU (CUDA) parallel implementations for the solution of an industrial-sized unstructured mesh consisting of about 1.5 million edges. Results show the significance of choosing the correct partition and thread-block configuration, the factors limiting the GPU performance and insights into optimizations for improved performance. Michael B. Giles, Gihan R. Mudalige, Z. Sharif, Graham R. Markall, Paul H. J. Kelly |
Comput. J. | 2 |
| 2012 | On the Acceleration of Wavefront Applications using Distributed Many-Core ArchitecturesabstractIn this paper we investigate the use of distributed graphics processing unit (GPU)-based architectures to accelerate pipelined wavefront applications—a ubiquitous class of parallel algorithms used for the solution of a number of scientific and engineering applications. Specifically, we employ a recently developed port of the LU solver (from the NAS Parallel Benchmark suite) to investigate the performance of these algorithms on high-performance computing solutions from NVIDIA (Tesla C1060 and C2050) as well as on traditional clusters (AMD/InfiniBand and IBM BlueGene/P). Benchmark results are presented for problem classes A to C and a recently developed performance model is used to provide projections for problem classes D and E, the latter of which represents a billion-cell problem. Our results demonstrate that while the theoretical performance of GPU solutions will far exceed those of many traditional technologies, the sustained application performance is currently comparable for scientific wavefront applications. Finally, a breakdown of the GPU solution is conducted, exposing PCIe overheads and decomposition constraints. A new k-blocking strategy is proposed to improve the future performance of this class of algorithm on GPU-based architectures. Simon J. Pennycook, Simon D. Hammond, Gihan R. Mudalige, Steven A. Wright 0001, Stephen A. Jarvis |
Comput. J. | 3 |
| 2009 | Predictive Simulation of HPC ApplicationsabstractThe architectures which support modern supercomputing machinery are as diverse today, as at any point during the last twenty years. The variety of processor core arrangements, threading strategies and the arrival of heterogeneous computation nodes are driving modern-day solutions to petaflop speeds. The increasing complexity of such systems, as well as codes written to take advantage of the new computational abilities, pose significant frustrations for existing techniques which aim to model and analyze the performance of such hardware and software. In this paper we demonstrate the use of post-execution analysis on trace-based profiles to support the construction of simulation-based models. This involves combining the runtime capture of call-graph information with computational timings, which in turn allows representative models of code behavior to be extracted. The main advantage of this technique is that it largely automates performance model development, a burden associated with existing techniques. We demonstrate the capabilities of our approach using both the NAS Parallel Benchmark suite and a real-world supercomputing benchmark developed by the United Kingdom Atomic Weapons Establishment. The resulting models, developed in less than two hours per code, have a good degree of predictive accuracy. We also show how one of these models can be used to explore the performance of the code on over 16,000 cores, demonstrating the scalability of our solution. Simon D. Hammond, J. A. Smith, Gihan R. Mudalige, Stephen A. Jarvis |
AINA | 3 |
| 2009 | Predictive analysis and optimisation of pipelined wavefront computationsabstractPipelined wavefront computations are a ubiquitous class of parallel algorithm used for the solution of a number of scientific and engineering applications. This paper investigates three optimisations to the generic pipelined wavefront algorithm, which are investigated through the use of predictive analytic models. The modelling of potential optimisations is supported by a recently developed reusable LogGP-based analytic performance model, which allows the speculative evaluation of each optimisation within the context of an industry-strength pipelined wavefront benchmark developed and maintained by the United Kingdom Atomic Weapons Establishment (AWE). The paper details the quantitative and qualitative benefits of: (1) parallelising computation blocks of the wavefront algorithm using OpenMP; (2) a novel restructuring/shifting of computation within the wavefront code and, (3) performing simultaneous multiple sweeps through the data grid. Gihan R. Mudalige, Simon D. Hammond, J. A. Smith, Stephen A. Jarvis |
IPDPS | 1 |
| 2008 | A plug-and-play model for evaluating wavefront computations on parallel architecturesabstractThis paper develops a plug-and-play reusable LogGP model that can be used to predict the runtime and scaling behavior of different MPI-based pipelined wavefront applications running on modern parallel platforms with multi-core nodes. A key new feature of the model is that it requires only a few simple input parameters to project performance for wavefront codes with different structure to the sweeps in each iteration as well as different behavior during each wavefront computation and/or between iterations. We apply the model to three key benchmark applications that are used in high performance computing procurement, illustrating that the model parameters yield insight into the key differences among the codes. We also develop new, simple and highly accurate models of MPI send, receive, and group communication primitives on the dual-core Cray XT system. We validate the reusable model applied to each benchmark on up to 8192 processors on the XT3/XT4. Results show excellent accuracy for all high performance application and platform configurations that we were able to measure. Finally we use the model to assess application and hardware configurations, develop new metrics for procurement and configuration, identify bottlenecks, and assess new application design modifications that, to our knowledge, have not previously been explored. Gihan R. Mudalige, Mary K. Vernon, Stephen A. Jarvis |
IPDPS | 1 |
| 2006 | Predictive Performance Analysis of a Parallel Pipelined Synchronous Wavefront Application for Commodity Processor Cluster SystemsabstractThis paper details the development and application of a model for predictive performance analysis of a pipelined synchronous wavefront application running on commodity processor cluster systems. The performance model builds on existing work (Cao et al.) by including extensions for modern commodity processor architectures. These extensions, including coarser hardware benchmarking, prove to be essential in countering the effects of modern superscalar processors (e.g. multiple operation pipelines and on-the-fly optimisations), complex memory hierarchies, and the impact of applying modern optimising compilers. The process of application modelling is also extended, combining static source code analysis with run-time profiling results for increased accuracy. The model is validated on several high performance SMP systems and the results show a high predictive accuracy (les 10% error). Additionally, the use of the performance model to speculate on the performance and scalability of this application on a hypothetical cluster with two different problem sizes is demonstrated. It is shown that such speculative techniques can be used to support system procurement, run-time verification and system maintenance and upgrading Gihan R. Mudalige, Stephen A. Jarvis, Daniel P. Spooner, Graham R. Nudd |
CLUSTER | 1 |