VLDB 2026 Research / reviewers in the wild / expert
István Z. Reguly
dblp:129/5523 · also István Zoltan Reguly
· DBLP profile ↗
24ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0002-4385-4204ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 5 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reduced and mixed precision turbulent flow simulations using explicit finite difference schemesabstractThe use of reduced and mixed precision computing has gained increasing attention in high-performance computing (HPC) as a means to improve computational efficiency, particularly on modern hardware architectures like GPUs. In this work, we explore the application of mixed precision arithmetic in compressible turbulent flow simulations using explicit finite difference schemes. We extend the OPS and OpenSBLI frameworks to support customizable precision levels, enabling fine-grained control over precision allocation for different computational tasks. Through a series of numerical experiments on the Taylor–Green vortex benchmark, we demonstrate that mixed precision strategies, such as half-single and single-double combinations, can offer significant performance gains without compromising numerical accuracy. However, pure half-precision computations result in unacceptable accuracy loss, underscoring the need for careful precision selection. Our results show that mixed precision configurations can reduce memory usage and communication overhead, leading to notable speedups, particularly on multi-CPU and multi-GPU systems. Bálint Siklósi, Pushpender K. Sharma, David J. Lusher, István Z. Reguly, Neil D. Sandham |
Future Gener. Comput. Syst. | 4 |
| 2025 | Performance and efficiency: A multi-generational benchmark of modern processors on bandwidth-bound HPC applications
Balázs Drávai, István Z. Reguly |
Future Gener. Comput. Syst. | 2 |
| 2025 | Smart epidemic control: A hybrid model blending ODEs and agent-based simulations for optimal, real-world intervention planningabstractOptimal intervention planning is a critical part of epidemiological control, which is difficult to attain in real life situations. Ordinary differential equation (ODE) models can be used to optimize control but the results can not be easily translated to interventions in highly complex real life environments. Agent-based methods on the other hand allow detailed modeling of the environment but optimization is precluded by the large number of parameters. Our goal was to combine the advantages of both approaches, i.e., to allow control optimization in complex environments. The epidemic control objectives are expressed as a time-dependent reference for the number of infected people. To track this reference, a model predictive controller (MPC) is designed with a compartmental ODE prediction model to compute the optimal level of stringency of interventions, which are later translated to specific actions such as mobility restriction, quarantine policy, masking rules, school closure. The effects of interventions on the transmission rate of the pathogen, and hence their stringency, are computed using PanSim, an agent-based epidemic simulator that contains a detailed model of the environment. The realism and practical applicability of the method is demonstrated by the wide range of discrete level measures that can be taken into account. Moreover, the change between measures applied during consecutive planning intervals is also minimized. We found that such a combined intervention planning strategy is able to efficiently control a COVID-19-like epidemic process, in terms of incidence, virulence, and infectiousness with surprisingly sparse (e.g. 21 day) intervention regimes. At the same time, the approach proved to be robust even in scenarios with significant model uncertainties, such as unknown transmission rate, uncertain time and probability constants. The high performance of the computation allows a large number of test cases to be run. The proposed computational framework can be reused for epidemic management of unexpected pandemic events and can be customized to the needs of any country. Péter Polcz, István Z. Reguly, Kalman Tornai, János Juhász, Sándor Pongor, Attila Csikász-Nagy, Gábor Szederkényi |
PLoS Comput. Biol. | 2 |
| 2023 | Communication-Avoiding Optimizations for Large-Scale Unstructured-Mesh Applications with OP2abstractIn this paper, we investigate data movement-reducing and communication-avoiding optimizations and their practicable implementation for large-scale unstructured-mesh applications. Utilizing the high-level abstraction of the OP2 DSL for the unstructured-mesh class of codes, we reason about techniques for reduced communications across a consecutive sequence of loops – a loop-chain. The careful trade-off with increased redundant computation in place of data movement is analyzed for distributed-memory parallelization. A new communication-avoiding (CA) back-end for OP2 is designed, codifying these techniques such that they can be applied automatically to any OP2 application. The back-end is extended to operate on a cluster of GPUs, integrating GPU-to-GPU communication with CUDA, in combination with MPI. The new CA back-end is applied automatically to two non-trivial applications, including the OP2 version of Rolls-Royce’s production CFD application, Hydra. Performance is investigated on both CPU and GPU clusters on representative problems of 8M and 24M node mesh sizes. Results demonstrate how for select configurations the new CA back-end provides between 30 – 65% runtime reductions for the loop-chains in these applications for the mesh sizes on both an HPE Cray EX system and an NVIDIA V100 GPU cluster. We model and examine the determinants and characteristics of a given unstructured-mesh loop-chain that can lead to performance benefits with CA techniques, providing insights into the general feasibility and profitability of using the optimizations for this class of applications. Suneth Dasantha Ekanayake, István Z. Reguly, Fabio Luporini, Gihan R. Mudalige |
ICPP | 2 |
| 2022 | Towards Virtual Certification of Gas Turbine Engines With Performance-Portable SimulationsabstractWe present the large-scale, computational fluid dy-namics (CFD) simulation of a full gas-turbine engine compressor, demonstrating capability towards overcoming current limitations for virtual certification of aero-engine design. The simulation is carried out through a performance portable code-base on multi-core/many-core HPC clusters with a CFD-to-CFD coupled execution, combining an industrial CFD solver linked using custom coupler software. The application innovates in its design for performance portability through the OP2 domain specific library for the CFD components, allowing the automatic generation of highly optimized platform-specific parallelizations for both multi-core (CPU) and many-core (GPU) clusters from a single high-level source. The code is used for the simulation of a 4.58B node, full-annulus 10-row production-grade test compressor (DLR's Rig250), using a coupled sliding-plane setup on the ARCHER2 and Cirrus supercomputers at EPCC. The OP2 generated multiple parallelizations, together with optimized coupler configurations on heterogeneous/hybrid settings achieve, for the first time, execution of 1 revolution in less than 6 hours on 512 nodes of ARCHER2 (65k cores), with a parallel scaling efficiency of over 80 % compared to a 107 node run. Results indicate a speed up of the CFD suite by an order of a magnitude (≈30 x) relative to current production capability. Benchmarking and performance modelling project a time-to-solution of less than 5 hours on a cluster of 488xNVIDIA V100 GPUs, about 3x-4 x speedup over CPU clusters. The work demonstrates a step-change towards achieving virtual certification of aircraft engines with the requisite fidelity and tractable time-to-solution that was previously out of reach under production settings. Gihan R. Mudalige, István Z. Reguly, Arun Prabhakar, Dario Amirante, Leigh Lapworth, Stephen A. Jarvis |
CLUSTER | 2 |
| 2022 | High throughput multidimensional tridiagonal system solvers on FPGAsabstractWe present a high performance tridiagonal solver library for Xilinx FPGAs optimized for multiple multi-dimensional systems common in real-world applications. An analytical performance model is developed and used to explore the design space and obtain rapid performance estimates that are over 85% accurate. This library achieves an order of magnitude better performance when solving large batches of systems than previous FPGA work. A detailed comparison with a current state-of-the-art GPU library for multi-dimensional tridiagonal systems on an Nvidia V100 GPU shows the FPGA achieving competitive or better runtime and significant energy savings of over 30%. Through this design, we learn lessons about the types of applications where FPGAs can challenge the current dominance of GPUs. Kamalakkannan Kamalavasan, Gihan R. Mudalige, István Z. Reguly, Suhaib A. Fahmy |
ICS | 3 |
| 2022 | Microsimulation based quantitative analysis of COVID-19 management strategiesabstractPandemic management requires reliable and efficient dynamical simulation to predict and control disease spreading. The COVID-19 (SARS-CoV-2) pandemic is mitigated by several non-pharmaceutical interventions, but it is hard to predict which of these are the most effective for a given population. We developed the computationally effective and scalable, agent-based microsimulation framework PanSim, allowing us to test control measures in multiple infection waves caused by the spread of a new virus variant in a city-sized societal environment using a unified framework fitted to realistic data. We show that vaccination strategies prioritising occupational risk groups minimise the number of infections but allow higher mortality while prioritising vulnerable groups minimises mortality but implies an increased infection rate. We also found that intensive vaccination along with non-pharmaceutical interventions can substantially suppress the spread of the virus, while low levels of vaccination, premature reopening may easily revert the epidemic to an uncontrolled state. Our analysis highlights that while vaccination protects the elderly from COVID-19, a large percentage of children will contract the virus, and we also show the benefits and limitations of various quarantine and testing scenarios. The uniquely detailed spatio-temporal resolution of PanSim allows the design and testing of complex, specifically targeted interventions with a large number of agents under dynamically changing conditions. István Z. Reguly, Dávid Csercsik, János Juhász, Kalman Tornai, Zsófia Bujtár, Gergely Horváth, Bence Keömley-Horváth, Tamás Kós, György Cserey, Kristóf Iván, Sándor Pongor, Gábor Szederkényi, Gergely Röst, Attila Csikász-Nagy |
PLoS Comput. Biol. | 1 |
| 2021 | Automatic Parallelisation of Sturctured Mesh Computations with SYCL
Gábor Dániel Balogh, István Z. Reguly |
CLUSTER | 2 |
| 2021 | Predictive Analysis of Large-Scale Coupled CFD Simulations with the CPX Mini-AppabstractAs the complexity of multi-physics simulations increases, there is a need for efficient flow of information between components. Discrete ‘coupler’ codes can abstract away this process, improving solver interoperability. One such multi-physics problem is modelling the high pressure compressor of turbofan engines, where instances of rotor/stator CFD simulations are coupled. Configuring couplers and allocating resources correctly can be challenging for such problems due to the sliding interfaces between codes. In this research, we present CPX, a mini-coupler designed to model the performance behaviour of a production coupler framework at Rolls-Royce plc., used for coupling rotor/stator simulations. CPX, the first mini-coupler framework of its kind, is combined with a CFD mini-app to predict the run-time and scaling behaviour of large scale coupled CFD simulations. We demonstrate high qualitative and quantitative predictive accuracy with a less than 17 % mean error. A performance model is developed to predict the ‘optimum’ configuration of resources, and is tested to show the high accuracy of these predictions. The model is also used to project the ‘optimum’ configuration for a 6 Billion cell test case, a problem size representative of current leading-edge production workloads, on a 100,000 core cluster and a 400 GPU cluster. Further testing reveals that the ‘optimum’ configuration is unstable if not set up correctly, and therefore a trade-off needs to be made with a marginally less-than-optimal setup to ensure stability. The work illustrates the significant utility of CPX to carry out such rapid design space and run-time setup exploration studies to obtain the best performance from production CFD coupled simulations. Archie Powell, K. Choudry, Arun Prabhakar, István Z. Reguly, Dario Amirante, Stephen A. Jarvis, Gihan R. Mudalige |
HiPC | 4 |
| 2021 | High-Level FPGA Accelerator Design for Structured-Mesh-Based Explicit Numerical SolversabstractThis paper presents a workflow for synthesizing near-optimal FPGA implementations of structured-mesh based stencil applications for explicit solvers. It leverages key characteristics of the application class and its computation-communication pattern and the architectural capabilities of the FPGA to accelerate solvers for high-performance computing applications. Key new features of the workflow are (1) the unification of standard state-of-the-art techniques with a number of high-gain optimizations such as batching and spatial blocking/tiling, motivated by increasing throughput for real-world workloads and (2) the development and use of a predictive analytical model to explore the design space, and obtain resource and performance estimates. Three representative applications are implemented using the design workflow on a Xilinx Alveo U280 FPGA, demonstrating near-optimal performance and over 85% predictive model accuracy. These are compared with equivalent highly-optimized implementations of the same applications on modern HPC-grade GPUs (Nvidia V100), analyzing time to solution, bandwidth, and energy consumption. Performance results indicate comparable runtimes with the V100 GPU, with over 2× energy savings for the largest non-trivial application on the FPGA. Our investigation shows the challenges of achieving high performance on current generation FPGAs compared to traditional architectures. We discuss determinants for a given stencil code to be amenable to FPGA implementation, providing insights into the feasibility and profitability of a design and its resulting performance. Kamalakkannan Kamalavasan, Gihan R. Mudalige, István Z. Reguly, Suhaib A. Fahmy |
IPDPS | 3 |
| 2020 | Automatic parallel implementations of adjoint codes for structured mesh applicationsabstractAlgorithmic Differentiation (AD) shown to be an essential tool to get sensitivity information for va in multiple areas of science such as Computational Fluid Dynamics (CFD) applications or finance. Yet there is no sufficient tool to ease the cost of providing performance portable AD codes, especially for modern hardware like GPU clusters. This paper sketches our plans and progress so far to extend the OPS framework with an adjoint tape (storage for descriptors of intermediate steps and intermediate states of variables) and shows preliminary performance results on CPU nodes. The OPS (Oxford Parallel library for Structured mesh solvers) has shown good performance and scaling on a wide range of HPC architectures. Our work aims to exploit the benefits of OPS to provide performance portable adjoint implementations for future structured mesh stencil applications using OPS with minimal modifications. Gábor Dániel Balogh, István Z. Reguly |
CCGRID | 2 |
| 2020 | Bitwise Reproducible task execution on unstructured mesh applicationsabstractMany mesh applications use floating point arithmetic which do not necessarily hold the associative laws of algebra. This could cause the application to become unreproducible. In this paper we present some work on generating a method for unstructured mesh applications to provide bitwise reproducibility between separate runs, even if they are started with different number of MPI processes. We implement our work in the OP2 domain-specific library, which provides an API that abstracts the solution of unstructured mesh computations. We carry out a performance analysis of our method applied on two applications: a simple airfoil application, and a more complex Aero application which uses a finite element method and a conjugate-gradient algorithm. We show a 2.37×to 1.49× slowdown on this applications as a price for full bitwise reproducibility. Bálint Siklósi, István Z. Reguly, Gihan R. Mudalige |
CCGRID | 2 |
| 2019 | PPCU Sam: Open-source face recognition frameworkabstractIn recent years by the popularization of AI, an increasing number of enterprises deployed machine learning algorithms in real life settings. This trend shed light on leaking spots of the Deep Learning bubble, namely the catastrophic decrease in quality when the distribution of the test data shifts from the training data. It is of utmost importance that we treat the promises of novel algorithms with caution and discourage reporting near perfect experimental results by fine-tuning on fixed test sets and finding metrics that hide weak points of the proposed methods. To support the wider acceptance of computer vision solutions we share our findings through a case-study in which we built a face-recognition system from scratch using consumer grade devices only, collected a database of 100k images from 150 subjects and carried out extensive validation of the most prominent approaches in single-frame face recognition literature. We show that the reported worst-case score, 74.3% true-positive ratio drops below 46.8% on real data. To overcome this barrier, after careful error analysis of the single-frame baselines we propose a low complexity solution to cover the failure cases of the single-frame recognition methods which yields an increased stability in multi-frame recognition during test time. We validate the effectiveness of the proposal by an extensive survey among our users which evaluates to 89.5% true-positive ratio. Botos Csaba, Hakkel Tamás, András Horváth, Andras Olah, István Z. Reguly |
KES | 5 |
| 2019 | Large-scale performance of a DSL-based multi-block structured-mesh application for Direct Numerical Simulation
Gihan R. Mudalige, István Z. Reguly, Satya P. Jammy, Christian T. Jacobs, Michael B. Giles, Neil D. Sandham |
J. Parallel Distributed Comput. | 2 |
| 2019 | Improving resilience of scientific software through a domain-specific approach
István Z. Reguly, Gihan R. Mudalige, Michael B. Giles, S. Maheswaran 0003 |
J. Parallel Distributed Comput. | 1 |
| 2019 | Locality optimized unstructured mesh algorithms on GPUs
András Attila Sulyok, Gábor Dániel Balogh, István Z. Reguly, Gihan R. Mudalige |
J. Parallel Distributed Comput. | 3 |
| 2018 | Loop Tiling in Large-Scale Stencil Codes at Run-Time with OPSabstractThe key common bottleneck in most stencil codes is data movement, and prior research has shown that improving data locality through optimisations that optimise across loops do particularly well. However, in many large PDE applications it is not possible to apply such optimisations through compilers because there are many options, execution paths and data per grid point, many dependent on run-time parameters, and the code is distributed across different compilation units. In this paper, we adapt the data locality improving optimisation called tiling for use in large OPS applications both in shared-memory and distributed-memory systems, relying on run-time analysis and delayed execution. We evaluate our approach on a number of applications, observing speedups of 2× on the Cloverleaf 2D/3D proxy applications, which contain 83(2D)/141(3D) loops, 3.5× on the linear solver TeaLeaf, and 1.7× on the compressible Navier-Stokes solver OpenSBLI. We demonstrate strong and weak scalability on up to 4608 cores of CINECA's Marconi supercomputer. We also evaluate our algorithms on Intel's Knights Landing, demonstrating maintained throughput as the problem size grows beyond 16GB, and we do scaling studies up to 8704 cores. The approach is generally applicable to any stencil DSL that provides per loop nest data access information. István Z. Reguly, Gihan R. Mudalige, Michael B. Giles |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2017 | Achieving Performance Portability for a Heat Conduction Solver Mini-Application on Modern Multi-core SystemsabstractModernizing production-grade, often legacy applications to take advantage of modern multi-core and many-core architectures can be a difficult and costly undertaking. This is especially true currently, as it is unclear which architectures will dominate future systems. The complexity of these codes can mean that parallelisation for a given architecture requires significant re-engineering. One way to assess the benefit of such an exercise would be to use mini-applications that are representative of the legacy programs.In this paper, we investigate different implementations of TeaLeaf, a mini-application from the Mantevo suite that solves the linear heat conduction equation. TeaLeaf has been ported to use many parallel programming models, including OpenMP, CUDA and MPI among others. It has also been re-engineered to use the OPS embedded DSL and template libraries Kokkos and RAJA. We use these different implementations to assess the performance portability of each technique on modern multi-core systems.While manually parallelising the application targeting and optimizing for each platform gives the best performance, this has the obvious disadvantage that it requires the creation of different versions for each and every platform of interest. Frameworks such as OPS, Kokkos and RAJA can produce executables of the program automatically that achieve comparable portability. Based on a recently developed performance portability metric, our results show that OPS and RAJA achieve an application performance portability score of 71% and 77% respectively for this application. Richard O. Kirk, Gihan R. Mudalige, István Z. Reguly, Steven A. Wright 0001, Matt Martineau, Stephen A. Jarvis |
CLUSTER | 3 |
| 2016 | Vectorizing unstructured mesh computations for many-core architecturesabstractSummary Achieving optimal performance on the latest multi‐core and many‐core architectures increasingly depends on making efficient use of the hardware's vector units. This paper presents results on achieving high performance through vectorization on CPUs and the Xeon‐Phi on a key class of irregular applications: unstructured mesh computations. Using single instruction multiple thread (SIMT) and single instruction multiple data (SIMD) programming models, we show how unstructured mesh computations map to OpenCL or vector intrinsics through the use of code generation techniques in the OP2 Domain Specific Library and explore how irregular memory accesses and race conditions can be organized on different hardware. We benchmark Intel Xeon CPUs and the Xeon‐Phi, using a tsunami simulation and a representative CFD benchmark. Results are compared with previous work on CPUs and NVIDIA GPUs to provide a comparison of achievable performance on current many‐core systems. We show that auto‐vectorization and the OpenCL SIMT model do not map efficiently to CPU vector units because of vectorization issues and threading overheads. In contrast, using SIMD vector intrinsics imposes some restrictions and requires more involved programming techniques but results in efficient code and near‐optimal performance, two times faster than non‐vectorized code. We observe that the Xeon‐Phi does not provide good performance for these applications but is still comparable with a pair of mid‐range Xeon chips. Copyright © 2015 John Wiley & Sons, Ltd. István Z. Reguly, Endre László, Gihan R. Mudalige, Michael B. Giles |
Concurr. Comput. Pract. Exp. | 1 |
| 2016 | Acceleration of a Full-Scale Industrial CFD Application with OP2abstractHydra is a full-scale industrial CFD application used for the design of turbomachinery at Rolls Royce plc., capable of performing complex simulations over highly detailed unstructured mesh geometries. Hydra presents major challenges in data organization and movement that need to be overcome for continued high performance on emerging platforms. We present research in achieving this goal through the OP2 domain-specific high-level framework, demonstrating the viability of such a high-level programming approach. OP2 targets the domain of unstructured mesh problems and enables execution on a range of back-end hardware platforms. We chart the conversion of Hydra to OP2, and map out the key difficulties encountered in the process. Specifically we show how different parallel implementations can be achieved with an active library framework, even for a highly complicated industrial application and how different optimizations targeting contrasting parallel architectures can be applied to the whole application, seamlessly, reducing developer effort and increasing code longevity. Performance results demonstrate that not only the same runtime performance as that of the hand-tuned original code could be achieved, but it can be significantly improved on conventional processor systems, and many-core systems. Our results provide evidence of how high-level frameworks such as OP2 enable portability across a wide range of contrasting platforms and their significant utility in achieving high performance without the intervention of the application programmer. István Z. Reguly, Gihan R. Mudalige, Carlo Bertolli, Michael B. Giles, Adam Betts, Paul H. J. Kelly, David Radford |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2015 | Design and Development of Domain Specific Active Libraries with Proxy ApplicationsabstractRepresentative applications are versatile tools to evaluate new programming approaches, techniques and optimisations as a way to ensure continued high performance on future computing architectures. They make experimentation much easier before adopting changes/insights into the large scientific codes. In this paper we demonstrate the important role played by representative/proxy applications in designing and developing two high-level programming approaches: namely the OP2 and OPS domain specific (active) libraries. OP2 and OPS utilizes code generation techniques to produce automatic parallelisations from a high-level abstract problem declaration. The strategy delivers significant developer productivity to the domain scientist, while at the same time allowing computational experts to adopt the latest programming models and hardware-specific optimisations into the library and code generation tools to achieve near optimal performance. We show how representative applications have been a cornerstone in the development of OP2 and OPS and chart our experiences. In particular, we demonstrate how the range of hand-tuned optimized parallelisations of the CloverLeaf hydrodynamics mini-app allowed us to gain clear evidence that the OPS based code generated parallelisations were indeed as optimal as the hand-tuned versions. Additionally, with the use of a representative application from the CFD domain we demonstrate how the optimisations discovered and applied to proxy apps are indeed directly transferable to a large-scale industrial application at Rolls Royce plc. These results provide significant evidence into the utility of representative applications to improve productivity, enable performance portability and ultimately future-proof scientific applications. István Z. Reguly, Gihan R. Mudalige, Michael B. Giles |
CLUSTER | 1 |
| 2015 | Analysis of parallel processor architectures for the solution of the Black-Scholes PDEabstractCommon parallel computer microarchitectures offer a wide variety of solutions to implement numerical algorithms. The efficiency of different algorithms applied to the same problem vary with the underlying architecture which can be a multi-core CPU, many-core GPU, Intel's MIC (Many Integrated Core) or FPGA architecture. Significant differences between these architectures exist in the ISA (Instruction Set Architecture) and the way the compute flow is executed. The way parallelism is expressed changes with the ISA, thread management and customization available on the device. These differences pose restrictions to the implementable algorithms. The aim of the work is to analyze the efficiency of the algorithms through the architectural differences. The problem at hand is the one-factor Black-Scholes option pricing equation which is a parabolic PDE solved with explicit and implicit time-marching algorithms. In the implicit solution a scalar tridiagonal system of equations needs to be solved. The possible CPU, GPU implementations along with novel FPGA solutions with HLS (High Level Synthesis) will be shown. Performance is also analyzed and remarks on efficiency are made. Endre László, Zoltán Nagy 0001, Michael B. Giles, István Z. Reguly, Jeremy Appleyard, Péter Szolgay |
ISCAS | 4 |
| 2013 | Designing OP2 for GPU architectures
Michael B. Giles, Gihan R. Mudalige, B. Spencer, Carlo Bertolli, István Z. Reguly |
J. Parallel Distributed Comput. | 5 |
| 2013 | Design and initial performance of a high-level unstructured mesh framework on heterogeneous parallel systems
Gihan R. Mudalige, Michael B. Giles, Jeyan Thiyagalingam, István Z. Reguly, Carlo Bertolli, Paul H. J. Kelly, Anne E. Trefethen |
Parallel Comput. | 4 |