VLDB 2026 Research / reviewers in the wild / expert
Nick Brown 0002
dblp:123/6117-2
· DBLP profile ↗
24ranked-venue papers
12as first author
15since 2021 · last 2026
0000-0003-2925-7275ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 12 first-author · 14 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An MLIR Lowering Pipeline for Stencils at Wafer-ScaleabstractThe Cerebras Wafer-Scale Engine (WSE) delivers performance at an unprecedented scale of over 900,000 compute units, all connected via a single-wafer on-chip interconnect. Initially designed for AI, the WSE architecture is also well-suited for High Performance Computing (HPC). However, its distributed asynchronous programming model diverges significantly from the simple sequential or bulk-synchronous programs that one would typically derive for a given mathematical program description. Targeting the WSE requires a bespoke re-implementation when porting existing code. The absence of WSE support in compilers such as MLIR, meant that there was little hope for automating this process. Nicolai Stawinoga, David Katz, Anton Lydike, Justs Zarins, Nick Brown 0002, George Bisbas, Tobias Grosser |
ASPLOS (2) | 5 |
| 2026 | Lifting to Tensors when Compiling Scientific Computing Workloads for AI Engines
Nick Brown 0002, Gabriel Rodriguez-Canal |
CCGrid | 1 |
| 2026 | Towards Compiler-Driven Dynamic Partial Reconfiguration with MLIRabstractHigh-Level Synthesis (HLS) has democratised Field-Programmable Gate Array (FPGA) programming, yet Dynamic Partial Reconfiguration (DPR)—which enables runtime logic swapping for adaptive or oversized workloads—remains manual and expert-only. HiPR [1] adds limited compiler support but restricts modules to one-to-one region mappings without runtime management. MLIR-DPR introduces: (i) a dpr dialect in the Multi-Level Intermediate Representation (MLIR) infrastructure [2] for identifying mutually exclusive regions; (ii) automated interface synthesis, floorplanning, and multi-threaded scheduler generation; and (iii) demonstrated Software-Defined Radio (SDR), Design-Space Exploration (DSE), and virtual-area applications. Gabriel Rodriguez-Canal, Nick Brown 0002, Maurice Jamieson, Nicolas Bohm Agostini, Ankur Limaye, Vito Giovanni Castellana, Joseph B. Manzano, Antonino Tumeo |
FCCM | 2 |
| 2026 | Towards Scheduling of Pipelined Dataflow Graphs in MLIRabstractWe present an MLIR flow that partitions neural networks and schedules them as software-driven macro-dataflow pipelines for low-latency streaming on CPU–FPGA SoCs. A new dataflow dialect and token-based scheduler pipeline even cyclic graphs with external memory, overcoming HLS limits. On an AlphaData ADM-PA101 (Versal VM1802) we demonstrate low-latency streaming; to our knowledge this is the first HLS flow to pipeline cyclic NN graphs. Gabriel Rodriguez-Canal, Nicolas Bohm Agostini, Ankur Limaye, Vito Giovanni Castellana, Joseph B. Manzano, Antonino Tumeo, Maurice Jamieson, Nick Brown 0002 |
FPGA | 8 |
| 2025 | Seamless Acceleration of Fortran Intrinsics via AMD AI Engines
Nick Brown 0002, Gabriel Rodriguez-Canal |
FPGA | 1 |
| 2024 | A shared compilation stack for distributed-memory parallelism in stencil DSLsabstractDomain Specific Languages (DSLs) increase programmer productivity and provide high performance. Their targeted abstractions allow scientists to express problems at a high level, providing rich details that optimizing compilers can exploit to target current- and next-generation supercomputers. The convenience and performance of DSLs come with significant development and maintenance costs. The siloed design of DSL compilers and the resulting inability to benefit from shared infrastructure cause uncertainties around longevity and the adoption of DSLs at scale. By tailoring the broadly-adopted MLIR compiler framework to HPC, we bring the same synergies that the machine learning community already exploits across their DSLs (e.g. Tensorflow, PyTorch) to the finite-difference stencil HPC community. We introduce new HPC-specific abstractions for message passing targeting distributed stencil computations. We demonstrate the sharing of common components across three distinct HPC stencil-DSL compilers: Devito, PSyclone, and the Open Earth Compiler, showing that our framework generates high-performance executables based upon a shared compiler ecosystem. George Bisbas, Anton Lydike, Emilien Bauer, Nick Brown 0002, Mathieu Fehr, Lawrence Mitchell, Gabriel Rodriguez-Canal, Maurice Jamieson, Paul H. J. Kelly, Michel Steuwer, Tobias Grosser |
ASPLOS (3) | 4 |
| 2024 | Evaluating Versal AI Engines for Option Price Discovery in Market Risk AnalysisabstractWhilst Field-Programmable Gate Arrays (FPGAs) have been popular in accelerating high-frequency financial workload for many years, their application in quantitative finance, the utilisation of mathematical models to analyse financial markets and securities, is less mature. Nevertheless, recent work has demonstrated the benefits that FPGAs can deliver to quantitative workloads, and in this paper, we study whether the Versal ACAP and its AI Engines (AIEs) can also deliver improved performance. We focus specifically on the industry standard Strategic Technology Analysis Center's (STAC) derivatives risk analysis benchmark STAC-A2. Porting a purely FPGA-based accelerator STAC-A2 inspired market risk (SIMR) benchmark to the Versal ACAP device by combining Programmable Logic (PL) and AIEs, we explore the development approach and techniques, before comparing performance across PL and AIEs. Ultimately, we found that our AIE approach is slower than a highly optimised existing PL-only version due to limits on both the AIE and PL that we explore and describe. Mark Klaisoongnoen, Nick Brown 0002, Timothy Dykes, Jessica R. Jones, Utz-Uwe Haus |
FPGA | 2 |
| 2024 | Predicting accurate batch queue wait times on production supercomputers by combining machine learning techniquesabstractAbstract The ability to accurately predict when a job on a supercomputer will leave the queue and start to run is not only beneficial for providing insights to users, but can also help enable non‐traditional HPC workloads that are not necessarily suited to the batch queue style‐approach that is ubiquitous on production HPC machines. However there are numerous challenges in achieving such a prediction with high accuracy, not least because the queue's state can change rapidly and depend upon many factors. In this work, we explore a novel machine learning approach for predicting queue wait times, hypothesising that such a model can capture the complex behavior resulting from the queue policy and other interactions to generate accurate job start times. For ARCHER2 (HPE Cray EX), Cirrus (HPE 8600), and 4‐cabinet (HPE Cray EX) we explore how different machine learning approaches and techniques improve the accuracy of our predictions, comparing against the estimation generated by Slurm. By combining categorization and regression models, we demonstrate that our approach delivers the most accurate predictions across our machines of interest, with the result of this work being the ability to predict job start times within 1 min of the actual start time for around 65% of jobs on ARCHER2 and 4‐cabinet, and 76% of jobs on Cirrus. When compared against what Slurm can deliver, via the backfill plugin, this represents around 3.8 times better accuracy on ARCHER2 and 18 times better for Cirrus. Furthermore our approach can accurately predicting the start time for three quarters of all job within 10 min of the actual start time on ARCHER2 and 4‐cabinet, and for 90% of jobs on Cirrus. Whilst the initial driver of this work was to better facilitate non‐traditional, interactive and urgent, workloads on HPC machines, the insights gained can also be used to provide wider benefits to users, enrich existing batch queue systems, and inform supercomputing center policy also. Nick Brown 0002, Gordon Gibb, Evgenij Belikov, Rupert W. Nash |
Concurr. Comput. Pract. Exp. | 1 |
| 2023 | Exploring the Versal AI Engines for Accelerating Stencil-based Atmospheric Advection SimulationabstractAMD Xilinx's new Versal Adaptive Compute Acceleration Platform (ACAP) is an FPGA architecture combining reconfigurable fabric with other on-chip hardened compute resources. AI engines are one of these and, by operating in a highly vectorized manner, they provide significant raw compute that is potentially beneficial for a range of workloads including HPC simulation. However, this technology is still early-on, and as yet unproven for accelerating HPC codes, with a lack of benchmarking and best practice. Nick Brown 0002 |
FPGA | 1 |
| 2023 | Fortran High-Level Synthesis: Reducing the Barriers to Accelerating HPC Codes on FPGAsabstractIn recent years the use of FPGAs to accelerate scientific applications has grown, with numerous applications demonstrating the benefit of FPGAs for high performance workloads. However, whilst High Level Synthesis (HLS) has significantly lowered the barrier to entry in programming FPGAs by enabling programmers to use C++, a major challenge is that most often these codes are not originally written in C++. Instead, Fortran is the lingua franca of scientific computing and-so it requires a complex and time consuming initial step to convert into C++ even before considering the FPGA. In this paper we describe work enabling Fortran for AMD Xilinx FPGAs by connecting the LLVM Flang front end to AMD Xilinx's LLVM back end. This enables programmers to use Fortran as a first-class language for programming FPGAs, and as we demonstrate enjoy all the tuning and optimisation opportunities that HLS C++ provides. Furthermore, we demonstrate that certain language features of Fortran make it especially beneficial for programming FPGAs compared to C++. The result of this work is a lowering of the barrier to entry in using FPGAs for scientific computing, enabling programmers to leverage their existing codebase and language of choice on the FPGA directly. Gabriel Rodriguez-Canal, Nick Brown 0002, Timothy Dykes, Jessica R. Jones, Utz-Uwe Haus |
FPL | 2 |
| 2023 | Task-based preemptive scheduling on FPGAs leveraging partial reconfigurationabstractSummary Field‐programmable gate arrays (FPGAs) are an attractive type of accelerator for all‐purpose high performance computing computing systems due to the possibility of deploying tailored hardware on demand. However, the common tools for programming and operating FPGAs are still complex to use, especially in scenarios where diverse types of tasks should be dynamically executed. In this work, we present a programming abstraction with a simple interface that internally leverages high‐level synthesis, dynamic partial reconfiguration and synchronization mechanisms to use an FPGA as a multi‐tasking server with preemptive scheduling and priority queues. This leads to an improved use of the FPGA resources, allowing the execution of several different kernels concurrently and deploying the most urgent ones as fast as possible. The results of our experimental study show that our approach incurs only a 10 5% overhead in the worst case when using two reconfigurable regions, whilst providing a significant performance improvement of at least 24 21% over the traditional full reconfiguration approach. Gabriel Rodriguez-Canal, Nick Brown 0002, Yuri Torres, Arturo González-Escribano |
Concurr. Comput. Pract. Exp. | 2 |
| 2022 | A programming model for developing Application Specific Dataflow Machines on FPGAsabstractFPGAs have become popular in many fields but are yet to gain wide acceptance in High Performance Computing (HPC) for accelerating scientific or engineering simulations. Whilst there are numerous on-going activities exploring the role of FPGAs for such workloads, often using HLS which enables programming in C or C++, significant challenges remain for scientific software developers to achieve performance with reconfigurable architectures. The underlying issue is that HLS presents a Von Neumann based programming model which is inappropriate for FPGAs, resulting in a significant disconnect between the semantics of existing HLS-based languages, and how experienced FPGA programmers must write their dataflow codes to best exploit the hardware. Nick Brown 0002 |
FCCM | 1 |
| 2021 | Compact native code generation for dynamic languages on micro-core architecturesabstractMicro-core architectures combine many simple, low memory, low power-consuming CPU cores onto a single chip. Potentially providing significant performance and low power consumption, this technology is not only of great interest in embedded, edge, and IoT uses, but also potentially as accelerators for data-center workloads. Due to the restricted nature of such CPUs, these architectures have traditionally been challenging to program, not least due to the very constrained amounts of memory (often around 32KB) and idiosyncrasies of the technology. However, more recently, dynamic languages such as Python have been ported to a number of micro-cores, but these are often delivered as interpreters which have an associated performance limitation. Maurice Jamieson, Nick Brown 0002 |
CC | 2 |
| 2021 | Accelerating advection for atmospheric modelling on Xilinx and Intel FPGAsabstractReconfigurable architectures, such as FPGAs, execute code at the electronics level, avoiding assumptions imposed by the general purpose black-box micro-architectures of CPUs and GPUs. Such tailored execution can result in increased performance and power efficiency, and as the HPC community moves towards exascale an important question is the role these hardware technologies can play in future supercomputers.In this paper we explore the porting of the PW advection kernel, an important code component used in a variety of atmospheric simulations and accounting for around 40% of the runtime of the popular Met Office NERC Cloud model (MONC). Building upon previous work which ported this kernel to an older generation of Xilinx FPGA, we target latest generation Xilinx Alveo U280 and Intel Stratix 10 FPGAs. Designing around the abstraction of an Application Specific Dataflow Machine (ASDM), we develop a design which is performance portable between vendors and explore implementation differences between the tool chains and compare kernel performance between FPGA hardware. This is followed by a more general performance comparison, scaling up the number of kernels on the Xilinx Alveo and Intel Stratix 10, against a 24 core Xeon Platinum Cascade Lake CPU and NVIDIA Tesla V100 GPU. When overlapping the transfer of data to and from the boards with compute, the FPGA solutions considerably outperform the CPU and, whilst falling short of the GPU in terms of performance, demonstrate power usage benefits, with the Alveo being especially power efficient. The result of this work is a comparison and set of design techniques that apply both to this specific atmospheric advection kernel on Xilinx and Intel FPGAs, and that are also of interest more widely when looking to accelerate HPC codes on a variety of reconfigurable architectures. Nick Brown 0002 |
CLUSTER | 1 |
| 2021 | Optimisation of an FPGA Credit Default Swap engine by embracing dataflow techniquesabstractQuantitative finance is the use of mathematical models to analyse financial markets and securities. Typically requiring significant amounts of computation, an important question is the role that novel architectures can play in accelerating these models in the future on HPC machines. In this paper we explore the optimisation of an existing, open source, FPGA based Credit Default Swap (CDS) engine using High Level Synthesis (HLS). Developed by Xilinx, and part of their open source Vitis libraries, the implementation of this engine currently favours flexibility and ease of integration over performance.We explore redesigning the engine to fully embrace the dataflow approach, ultimately resulting in an engine which is around eight times faster on an Alveo U280 FPGA than the original Xilinx library version. We then compare five of our engines on the U280 against a 24-core Xeon Platinum Cascade Lake CPU, outperforming the CPU by around 1.55 times, with the FPGA consuming 4.7 times less power and delivering around seven times the power efficiency of the CPU. Nick Brown 0002, Mark Klaisoongnoen, Oliver Thomson Brown |
CLUSTER | 1 |
| 2020 | Investigating Applications on the A64FXabstractThe A64FX processor from Fujitsu, being designed for computational simulation and machine learning applications, has the potential for unprecedented performance in HPC systems. In this paper, we evaluate the A64FX by benchmarking against a range of production HPC platforms that cover a number of processor technologies. We investigate the performance of complex scientific applications across multiple nodes, as well as single node and mini-kernel benchmarks. This paper finds that the performance of the A64FX processor across our chosen benchmarks often significantly exceeds other platforms, even without specific application optimisations for the processor instruction set or hardware. However, this is not true for all the benchmarks we have undertaken. Furthermore, the specific configuration of applications can have an impact on the runtime and performance experienced. Adrian Jackson, Michèle Weiland, Nick Brown 0002, Andrew Turner, Mark Parsons 0001 |
CLUSTER | 3 |
| 2020 | Machine Learning for Gas and Oil ExplorationabstractDrilling boreholes for gas and oil extraction is an expensive process and profitability strongly depends on characteristics of the subsurface. As profitability is a key success factor, companies in the industry utilise well logs to explore the subsurface beforehand. These well logs contain various characteristics of the rock around the borehole, which allow petrophysicists to determine the expected amount of contained hydrocarbon. However, these logs are often incomplete and, as a consequence, the subsequent analyses cannot exploit the full potential of the well logs. In this paper we demonstrate that Machine Learning can be applied to fill in the gaps and estimate missing values. We investigate how the amount of training data influences the accuracy of prediction and how to best design regression models (Gradient Boosting and neural network) to obtain optimal results. We then explore the models' predictions both quantitatively, tracking the prediction error, and qualitatively, capturing the evolution of the measured and predicted values for a given property with depth. Combining the findings has enabled us to develop a predictive model that completes the well logs, increasing their quality and potential commercial value. Vito Alexander Nordloh, Anna Roubícková, Nick Brown 0002 |
ECAI | 3 |
| 2020 | Weighing Up the New Kid on the Block: Impressions of using Vitis for HPC Software DevelopmentabstractThe use of reconfigurable computing, and FPGAs in particular, has strong potential in the field of High Performance Computing (HPC). However the traditionally high barrier to entry when it comes to programming this technology has, until now, precluded widespread adoption. To popularise reconfigurable computing with communities such as HPC, Xilinx have recently released the first version of Vitis, a platform aimed at making the programming of FPGAs much more a question of software development rather than hardware design. However a key question is how well this technology fulfils the aim, and whether the tooling is mature enough such that software developers using FPGAs to accelerate their codes is now a more realistic proposition, or whether it simply increases the convenience for existing experts. To examine this question we use the Himeno benchmark as a vehicle for exploring the Vitis platform for building, executing and optimising HPC codes, describing the different steps and potential pitfalls of the technology. The outcome of this exploration is a demonstration that, whilst Vitis is an excellent step forwards and significantly lowers the barrier to entry in developing codes for FPGAs, it is not a silver bullet and an underlying understanding of dataflow style algorithmic design and appreciation of the architecture is still key to obtaining good performance on reconfigurable architectures. Nick Brown 0002 |
FPL | 1 |
| 2020 | Modelling the Earth's geomagnetic environment on Cray machines using PETSc and SLEPcabstractSummary The British Geological Survey's global geomagnetic model, Model of the Earth's Magnetic Environment (MEME), is an important tool for calculating the strength and direction of the Earth's magnetic field, which is continually in flux. While the ability to collect data from ground‐based observation sites and satellites has grown rapidly, the memory bound nature of the original code has proved a significant limitation on the size of the modelling problem required. In this paper, we describe work done replacing the bespoke, sequential, eigensolver with that of the PETSc/SLEPc package for solving the system of normal equations. Adopting PETSc/SLEPc also required fundamental changes in how we built and distributed the data structures, and as such, we describe an approach for building symmetric matrices that provides good load balance and avoids the need for close coordination between the processes or replication of work. We also study the memory bound nature of the code from an irregular memory accesses perspective and combine detailed profiling with software cache prefetching to significantly optimise this. Performance and scaling characteristics are explored on ARCHER, a Cray XC30, where we achieved a speed up for the solver of 294 times by replacing the model's bespoke approach with SLEPc. Nick Brown 0002, Brian Bainbridge, Ciarán Beggan, William Brown 0006, Brian Hamilton, Susan Macmillan |
Concurr. Comput. Pract. Exp. | 1 |
| 2020 | Machine learning on Crays to optimize petrophysical workflows in oil and gas explorationabstractSummary The oil and gas industry is awash with sub‐surface data, which is used to characterize the rock and fluid properties beneath the seabed. This drives commercial decision making and exploration, but the industry relies upon highly manual workflows when processing data. A question is whether this can be improved using machine learning, complementing the activities of petrophysicists searching for hydrocarbons. In this paper, we present work using supervised learning with the aim of decreasing the petrophysical interpretation time down from over 7 days to 7 minutes. We describe the use of mathematical models that have been trained using raw well log data, to complete each of the four stages of a petrophysical interpretation workflow, in addition to initial data cleaning. We explore how the predictions from these models compare against the interpretations of human petrophysicists, and numerous options and techniques that were used to optimize the models. The result of this work is the ability, for the first time, to use machine learning for the entire petrophysical workflow. Nick Brown 0002, Anna Roubícková, Ioanna Lampaki, Lucy MacGregor, Michelle Ellis, Paola Vera de Newton |
Concurr. Comput. Pract. Exp. | 1 |
| 2020 | High level programming abstractions for leveraging hierarchical memories with micro-core architectures
Maurice Jamieson, Nick Brown 0002 |
J. Parallel Distributed Comput. | 2 |
| 2019 | Leveraging MPI RMA to optimize halo-swapping communications in MONC on Cray machinesabstractSummary Remote Memory Access (RMA), also known as single‐sided communications, provides a way for reading and writing directly into the memory of other processes without having to issue explicit message passing style communication calls. Previous studies have concluded that MPI RMA can provide increased communication performance over traditional MPI Point to Point (P2P), but these are based on synthetic benchmarks rather than real‐world codes. In this work, we replace the existing non‐blocking P2P communication calls in the Met Office NERC Cloud model, a mature code for modeling the atmosphere, with MPI RMA. We describe our approach in detail and discuss the options taken for correctness and performance. Experiments are performed on ARCHER, a Cray XC30, and Cirrus, an SGI ICE machine. We demonstrate on ARCHER that, by using RMA, we can obtain between a 5% and 10% reduction in communication time at each timestep on up to 32768 cores, which over the entirety of a run (with many timesteps) results in a significant improvement in performance compared to P2P on the Cray. However, RMA is not a silver bullet, and there are challenges when integrating RMA calls into existing codes: important optimizations are necessary to achieve good performance and library support is not universally mature, as is the case on Cirrus. In this paper, we discuss, in the context of a real‐world code, the lessons learned converting P2P to RMA, explore performance and scaling challenges, and contrast alternative RMA synchronization approaches in detail. Nick Brown 0002, Michael R. Bareford, Michèle Weiland |
Concurr. Comput. Pract. Exp. | 1 |
| 2019 | iPregel: Vertex-centric programmability vs memory efficiency and performance, why choose?abstractThe vertex-centric programming model, designed to improve the programmability in graph processing application writing, has attracted great attention over the years. Multiple shared memory frameworks that have implemented the vertex-centric interface all expose a common tradeoff: programmability against memory efficiency and performance. Our approach consists in preserving vertex-centric programmability, while implementing optimisations missing from FemtoGraph, developing new ones and designing these so they are transparent to a user’s application code, hence not impacting programmability. We therefore implemented our own shared memory vertex-centric framework iPregel, relying on in-memory storage and synchronous execution. In this paper, we evaluate it against FemtoGraph, whose characteristics are identical, but also an asynchronous counterpart GraphChi and the vertex-subset-centric framework Ligra. Our experiments include three of the most popular vertex-centric benchmark applications over 4 real-world publicly accessible graphs, which cover all orders of magnitude between a million to a billion edges. We then measure the execution time and the peak memory usage. Finally, we evaluate the programmability of each framework by comparing it against the original Pregel, Google’s closed-source implementation that started the whole area of vertex-centric programming. Experiments demonstrate that iPregel, like FemtoGraph, does not sacrifice vertex-centric programmability for additional performance and memory efficiency optimisations, which contrasts with GraphChi and Ligra. Sacrificing vertex-centric programmability allowed the latter to benefit from substantial performance and memory efficiency gains. However, experiments demonstrate that iPregel is up to 2300 times faster than FemtoGraph, as well as generating a memory footprint up to 100 times smaller. These results greatly change the situation; Ligra and GraphChi are up to 17,000 and 700 times faster than FemtoGraph but, when comparing against iPregel, this maximum speed-up drops to 10. Furthermore, on PageRank, it is iPregel that proves to be the fastest overall. When it comes to memory efficiency, the same observation applies; Ligra and GraphChi are 100 and 50 times lighter than FemtoGraph, but iPregel nullifies these benefits: it provides the same memory efficiency as Ligra and even proves to be 3 to 6 times lighter than GraphChi on average. In other words, iPregel demonstrates that preserving vertex-centric programmability is not incompatible with a competitive performance and memory efficiency. Ludovic Anthony Richard Capelli, Zhenjiang Hu 0002, Timothy A. K. Zakian, Nick Brown 0002, J. Mark Bull |
Parallel Comput. | 4 |
| 2018 | In situ data analytics for highly scalable cloud modelling on Cray machinesabstractSummary MONC is a highly scalable modelling tool for the investigation of atmospheric flows, turbulence, and cloud microphysics. Typical simulations produce very large amounts of raw data, which must then be analysed for scientific investigation. For performance and scalability reasons, this analysis and subsequent writing to disk should be performed in situ on the data as it is generated; however, one does not wish to pause the computation whilst analysis is carried out. In this paper, we present the analytics approach of MONC, where cores of a node are shared between computation and data analytics. By asynchronously sending their data to an analytics core, the computational cores can run continuously without having to pause for data writing or analysis. We describe our IO server framework and analytics workflow, which is highly asynchronous, along with solutions to challenges that this approach raises and the performance implications of some common configuration choices. The result of this work is a highly scalable analytics approach, and we illustrate on up to 32 768 computational cores of a Cray XC30 that there is minimal performance impact on the runtime when enabling data analytics in MONC and also investigate the performance and suitability of our approach on the KNL. Nick Brown 0002, Michèle Weiland, Adrian Hill, Ben Shipway |
Concurr. Comput. Pract. Exp. | 1 |