Henri Calandra

dblp:04/2999 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
3since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2Artificial intelligence and machine learning · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2021 Tabu-Driven Quantum Neighborhood Samplers
Charles Moussa, Hao Wang 0025, Henri Calandra, Thomas Bäck, Vedran Dunjko
EvoCOP3
2021 Parallel and Distributed Task-Based Kirchhoff Seismic Pre-Stack Depth Migration Application
abstract
Since the middle of the 1990s, message passing libraries are the most used technology to implement parallel and distributed scientific applications. However, they may not be a solution efficient enough on exascale machines since scalability issues will appear due to the increase in computing resources. Task-based programming models can be used to avoid collective communications like reductions, broadcast, or gather by transforming them into multiple operations on tasks. Then, these operations can be scheduled by the programming scheduler to place the data and computations in a way that optimizes and reduces the data communications. These properties could help to solve some MPI and exascale computing challenges.The oil and gas applications could also benefit from task-based programming properties. We developed a simplified version of the Kirchhoff seismic pre-stack depth migration, a subsurface exploration application, to experiment with HPX, a task-based programming model as well and MPI and MPI+OpenMP. Then, we perform strong scaling and weak scaling experiments on Pangea, Total supercomputer. We also study the variation of the number of OpenMP threads per MPI process. We show that the current task-based programming model schedulers lack the capability to completely manage the memory used and are not efficient enough to reduce the data migrations.
Jérôme Gurhem, Henri Calandra, Serge G. Petiton
ISPDC2
2021 Practical Quantum Computing: Solving the Wave Equation Using a Quantum Approach
abstract
In the last few years, several quantum algorithms that try to address the problem of partial differential equation solving have been devised: on the one hand, “direct” quantum algorithms that aim at encoding the solution of the PDE by executing one large quantum circuit; on the other hand, variational algorithms that approximate the solution of the PDE by executing several small quantum circuits and making profit of classical optimisers. In this work, we propose an experimental study of the costs (in terms of gate number and execution time on a idealised hardware created from realistic gate data) associated with one of the “direct” quantum algorithm: the wave equation solver devised in [32]. We show that our implementation of the quantum wave equation solver agrees with the theoretical big-O complexity of the algorithm. We also explain in great detail the implementation steps and discuss some possibilities of improvements. Finally, our implementation proves experimentally that some PDE can be solved on a quantum computer, even if the direct quantum algorithm chosen will require error-corrected quantum chips, which are not believed to be available in the short-term.
Adrien Suau, Gabriel Staffelbach, Henri Calandra
ACM Trans. Quantum Comput.3
2017 One-Way Wave Equation Migration at Scale on GPUs Using Directive Based Programming
abstract
One-Way Wave Equation Migration (OWEM) is a depth migration algorithm used for seismic imaging. A parallel version of this algorithm is widely implemented using MPI. Heterogenous architectures that use GPUs have become popular in the Top 500 because of their performance/power ratio. In this paper, we discuss the methodology and code transformations used to port OWEM to GPUs using OpenACC, along with the code changes needed for scaling the application up to 18,400 GPUs (more than 98%) of the Titan leadership class supercomputer at Oak Ridget National Laboratory. For the individual OpenACC kernels, we achieved an average of 3X speedup on a test dataset using one GPU as compared with an 8-core Intel Sandy Bridge CPU. The application was then run at large scale on the Titan supercomputer achieving a peak of 1.2 petaflops using an average of 5.5 megawatts. After porting the application to GPUs, we discuss how we dealt with other challenges of running at scale such as the application becoming more I/O bound and prone to silent errors. We believe this work will serve as valuable proof that directive-based programming models are a viable option for scaling HPC applications to heterogenous architectures.
Kshitij Mehta, Maxime R. Hugues, Oscar R. Hernandez, David E. Bernholdt, Henri Calandra
IPDPS5
2013 Evaluation of Successive CPUs/APUs/GPUs Based on an OpenCL Finite Difference Stencil
abstract
The AMD APU (Accelerated Processing Unit) architecture, which combines CPU and GPU cores on the same die, is promising for GPU applications which performance is bottlenecked by the low PCI Express communication rate. However the first APU generations still have different CPU and GPU memory partitions. Currently, the APU integrated GPUs are also less powerful than discrete GPUs. In this paper we therefore investigate the interest of APUs for scientific computing by evaluating and comparing the performance of two successive AMD APUs (family codename Llano and Trinity), two successive discrete GPUs (chip codename Cayman and Tahiti) and one hexa-core AMD CPU. For this purpose, we rely on a 3D finite difference stencil, that is optimized and tuned in OpenCL. We detail the most interesting optimizations for each architecture and show very good performance in OpenCL: up to 500 Gflops on Tahiti. Finally, our results show that APU integrated GPUs outperform CPUs, and that integrated GPUs of upcoming APUs may match discrete GPUs for problems with high communication requirements.
Henri Calandra, Romain Dolbeau, Pierre Fortin 0001, Jean Luc Lamotte, Issam Said
PDP1
2012 Fast seismic modeling and reverse time migration on a graphics processing unit cluster
abstract
SUMMARY We designed a fast parallel simulator that solves the acoustic wave equation on a graphics processing unit (GPU) cluster. Solving the acoustic wave equation in an oil exploration industrial context aims at speeding up seismic modeling and reverse time migration (RTM). We considered a finite difference approach on a regular mesh, in both two‐dimensional and three‐dimensional cases. The acoustic wave equation is solved in a constant density or a variable density domain. All the computations were carried out in single precision (both in the CPU reference implementation and in the GPU implementation), because double precision was not required in our context. We used Compute Unified Device Architecture to take advantage of the GPU computational power. We studied different implementations and their impact on the application performance. The described application handles all the steps of seismic modeling and RTM and is used to solve real‐world problems in an industrial production context. We obtained a speedup of 16 for RTM and up to 43 for the modeling application over a sequential code running on general‐purpose CPUs. A CPU rack versus a GPU rack comparison was described and showed a 4.3 speedup. Copyright © 2011 John Wiley & Sons, Ltd.
Rached Abdelkhalek, Henri Calandra, Olivier Coulaud, Guillaume Latu, Jean Roman
Concurr. Comput. Pract. Exp.2