Hervé Yviquel

dblp:21/10574 · DBLP profile ↗
← Back
19ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0003-1214-3431ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author
YearPublicationVenuePosition
2026 Multi -FPGA streaming using OpenMP
abstract
The growth in the demand for high-performance and power-efficient applications has led to an increasing interest in FPGA-based acceleration. FPGAs have been applied to a wide range of applications. Still, programming them can be a complex task, requiring extensive knowledge of tools and libraries, especially in multi-FPGA architectures. As such, the desire for tools and frameworks to ease the burden and abstract the knowledge of using FPGAs has increased. OpenMP, an already dominant parallel programming model in HPC, has been shown to be a successful approach to program multi-FPGA architecture. This work is based on the OMPC-F framework, which leverages the capability of OpenMP to offload computation to FPGAs. Although OMPC-F abstracts FPGA handling and task distribution from the final user, it does not support streaming computation based on multi-FPGA architectures. Streaming is widely used in FPGA designs to create a pipeline of computation between FPGA kernels. This work uses FPGA kernel binary information to synthesize streams as OpenMP buffers, while adapting the OpenMP dependency system accordingly. The proposal was evaluated in an AMD/Xilinx multi-FPGA system and shows speedups of the order of 6.84x in 8 FPGAs, scaling well with the addition of more kernels and FPGAs to the architecture. Moreover, compared to the regular approach to developing FPGA applications based on MPI+XRT communication, the proposed approach reduces the programming effort by 59% according to various code analysis metrics, resulting in a small average overhead of 4.3% when compared to the MPI+XRT programming model.
Pedro Henrique Di Francia Rosso, Rémy Neveu, Nusrat Jahan Lisa, Lucas B. da Silva, Hervé Yviquel, Sandro Rigo, Vanderlei Bonato, Guido Araujo
J. Parallel Distributed Comput.5
2025 Scalable OpenMP Remote Offloading via Asynchronous MPI and Coroutine-Driven Communication
Jhonatan Cléto, Guilherme Valarini, Márcio Machado Pereira, Guido Araujo, Hervé Yviquel
Euro-Par (3)5
2025 A Distributed and Storage-Aware Approach to Large-Scale Cholesky Factorization
abstract
Cholesky factorization is a core operation in scientific computing, yet its scalability is often constrained by memory limitations when processing extremely large dense matrices. This work introduces an out-of-core Cholesky factorization algorithm for symmetric positive-definite matrices that integrates GPU acceleration, block-wise lossless compression, and parallel I/O to overcome these limitations. The approach leverages the OMPC runtime for asynchronous task scheduling and employs HDF5 to store the matrix on disk, taking advantage of Lustre’s parallel I/O capabilities in distributed environments. Tiles are decompressed just-in-time on the GPU, significantly reducing host memory usage, storage footprint, and end-to-end data movement overhead—from disk through the CPU to the GPU—without compromising numerical accuracy. Experimental results show that the proposed method scales across 8 GPU nodes, successfully factorizing matrices up to 3M × 3M. In comparison, SLATE could only handle sizes up to 700K × 700K, with the proposed algorithm achieving up to 41% higher throughput. These results demonstrate the algorithm’s scalability and competitiveness beyond memory-constrained in-core solutions, offering a practical path for enabling extreme-scale scientific applications.
Carla Cusihuallpa, Rodrigo Ceccato, Sandro Rigo, Guido Araujo, Hervé Yviquel
SBAC-PAD5
2025 Super-Stencil: A Memory-Efficient Superstep Wave Propagation Method for Seismic Imaging
abstract
Wave propagation is a fundamental component of seismic imaging, a technique crucial to oil and gas exploration. Traditional finite-difference (FD) methods advance the wave-field one time step at a time. While effective, these methods are memory-bound and require multiple high-end GPU nodes to achieve acceptable performance. To address this, the superstep wavefield propagation technique was introduced, grouping multiple time steps into a single large operator to increase computational intensity. However, it incurs a prohibitive memory overhead, requiring the storage of hundreds of large matrices for realistic problems. In this work, we introduce Super-Stencil, a novel symbolic formulation of superstep propagation that eliminates the need to store intermediate operators. This drastically reduces memory consumption—by about 1009x compared to the original superstep method for 20 time steps—shifting the computational bottleneck from memory to compute. Although Super-Stencil incurs up to 9.1× longer execution time than superstep and 337.3× longer than the FD baseline in a sequential setting, it unlocks a new dimension of parallelism through a Parallel-in-Time execution strategy. By enabling simultaneous time and space parallelism, Super-Stencil transforms wave propagation into a compute-bound kernel, opening the door to aggressive optimization on modern parallel architectures. This makes it a compelling alternative for next-generation seismic imaging workflows where scalability is paramount.
George Gigilas, Pedro S. Peixoto, Hermes Senger, Hervé Yviquel
SBAC-PAD4
2025 Profiler-Guided Execution of Recurrent OpenMP Task Graphs on Heterogeneous Clusters
abstract
Distributed task-based execution models are well-suited for parallelizing irregular applications across clusters. OpenMP Cluster (OMPC) extends the traditional OpenMP tasking model to support distributed memory systems, leveraging a HEFT-based scheduler to improve resource utilization. However, the efficiency of such a scheduler depends heavily on accurate estimates of task execution and communication costs – information that is often difficult to obtain reliably and efficiently. To address this limitation, we propose a novel scheduling framework that combines the recent taskgraph directive introduced in OpenMP 6.0 with partial online profiling of iterative applications. Our approach performs quasi-static scheduling by recording task graphs at runtime and selectively profiling representative iterations to estimate performance. This information is interpolated and fed back into the scheduler to enhance decision-making. We demonstrate that our framework can improve scheduling quality with minimal overhead, making it suitable for long-running or repetitive workloads commonly found in High-Performance Computing (HPC) applications. We achieve up to 20% speedup for the total application and 4× speedup for scheduling.
Rémy Neveu, Rodrigo Ceccato, Adrian Munera, Sara Royuela, José Monsalve Diaz, Hervé Yviquel
SBAC-PAD6
2024 Combining Compression and Prefetching to Improve Checkpointing for Inverse Seismic Problems in GPUs
Thiago Maltempi, Sandro Rigo, Márcio Machado Pereira, Hervé Yviquel, Jessé Costa, Guido Araujo
Euro-Par (3)4
2024 DeepWave: A Software Stack for Parallelizing Deep Learning Models Used in Geophysics
abstract
This paper introduces DeepWave, a novel software stack, and methodology designed to integrate generative artificial intelligence into traditional seismic surveying techniques, significantly enhancing the computational efficiency of geophysical exploration. By utilizing advanced machine learning frameworks such as JAX, FLAX, and ALPA, DeepWave employs a parallelization strategy for image-to-image translation networks, optimizing the seismic data interpretation process. DeepWave reduces the computational demands of intensive geophysical algorithms, such as Full-waveform Inversion (FWI), while maintaining the accuracy required for detailed subsurface analysis. This method enables faster and more efficient processing of large seismic datasets, providing deeper insights into the Earth’s subsurface structures with reduced computational resources. The results demonstrate a substantial improvement in processing speed and resource management, establishing a new geophysical research and exploration standard.
Allan Pinto, Gustavo Leite, Márcio Machado Pereira, Hervé Yviquel, Sandro Rigo, Guido Araujo
SBAC-PAD4
2024 Ion-molecule collision cross-section calculations using trajectory parallelization in distributed systems
Samuel Cajahuaringa, Leandro Zanotto, Sandro Rigo, Hervé Yviquel, Munir S. Skaf, Guido Araujo
J. Parallel Distributed Comput.4
2023 Source Matching and Rewriting for MLIR Using String-Based Automata
abstract
A typical compiler flow relies on a uni-directional sequence of translation/optimization steps that lower the program abstract representation, making it hard to preserve higher-level program information across each transformation step. On the other hand, modern ISA extensions and hardware accelerators can benefit from the compiler’s ability to detect and raise program idioms to acceleration instructions or optimized library calls. Although recent works based on Multi-Level IR (MLIR) have been proposed for code raising, they rely on specialized languages, compiler recompilation, or in-depth dialect knowledge. This article presents Source Matching and Rewriting (SMR), a user-oriented source-code-based approach for MLIR idiom matching and rewriting that does not require a compiler expert’s intervention. SMR uses a two-phase automaton-based DAG-matching algorithm inspired by early work on tree-pattern matching. First, the idiom Control-Dependency Graph (CDG) is matched against the program’s CDG to rule out code fragments that do not have a control-flow structure similar to the desired idiom. Second, candidate code fragments from the previous phase have their Data-Dependency Graphs (DDGs) constructed and matched against the idiom DDG. Experimental results show that SMR can effectively match idioms from Fortran (FIR) and C (CIL) programs while raising them as BLAS calls to improve performance. Additional experiments also show performance improvements when using SMR to enable code replacement in areas like approximate computing and hardware acceleration.
Vinícius Couto Espindola, Luciano G. Zago, Hervé Yviquel, Guido Araujo
ACM Trans. Archit. Code Optim.3
2022 Ion-Molecule Collision Cross-Section Simulation using Linked-cell and Trajectory Parallelization
abstract
Ion Mobility coupled to Mass Spectrometry (IMMS) has become a highly valued tool for structural characterization of biological samples. In IM-MS the protein under investigation is ionized and accelerated by an electric field into a drift tube where it collides against a buffer gas. The separation of the gas-phase ions is then measured through the differences in their rotationally averaged Collision Cross-Section (CCS) values. The utility of the measured CCS for structural characterization critically depends on the validation against its theoretical calculation, which relies on intensive molecular mechanics simulation. Increasing the performance of CCS simulation is thus a relevant computational-chemistry research problem. This work shows that the combination of a linked-cell based algorithm with parallelization techniques can considerably increase the performance of CCS simulation. Experimental results reveal speedups from$\sim \mathbf{10}\times$up to$\sim \mathbf{400}\times$and parallelization efficiency greater than 0.98 when compared to High Performance Collision Cross Section (HPCCS), an optimized solution for CCS simulation. This reduces the CCS computation time from hours to minutes for a large range of proteins, making the proposed method the most performant approach to this problem nowadays, to the best of our knowledge.
Samuel Cajahuaringa, Leandro Zanotto, Daniel L. Z. Caetano, Sandro Rigo, Hervé Yviquel, Munir S. Skaf, Guido Araujo
SBAC-PAD5
2021 Enabling OpenMP Task Parallelism on Multi-FPGAs
abstract
FPGA-based accelerators have received increasing attention recently. Nevertheless, the amount of resources available on even the most powerful FPGA is still not enough to speed up very large workloads. To achieve that, FPGAs need to be interconnected in a Multi-FPGA architecture. However, programming such architecture is a challenging endeavor. This paper extends the OpenMP task-based computation offloading model to enable several FPGAs to work as a single Multi-FPGA architecture. Experimental results, for a set of OpenMP stencil applications running on a Multi-FPGA platform, have shown close to linear speedups as the number of FPGAs and IP-cores per FPGA increases.
Ramon Nepomuceno, Renan Sterle, Guilherme Valarini, Márcio Machado Pereira, Hervé Yviquel, Guido Araujo
FCCM5
2020 OmpTracing: Easy Profiling of OpenMP Programs
abstract
One of the greatest challenges of modern computing is the development of software for parallel execution. To address such challenge, programmers use profiling tools to record relevant operations, like the communications that the different parts of an application carried out during its execution. Profilers can be used to analyze the execution of the application as they enable the programmer to check its performance hot spots and sources of overhead. This paper introduces the OmpTracing library, a lightweight tool that eases the task of profiling OpenMP based applications without the need to inject expensive profiling code into the program. OmpTracing leverages on OMPT, an application programming interface that provides an introspection mechanism of the OpenMP runtime, and that enables the programmer to capture execution details of the parallelized application while generating notifications about significant program events.
Vitoria Pinho, Hervé Yviquel, Márcio Machado Pereira, Guido Araujo
SBAC-PAD2
2018 Automatic Ray-Tracer Cloud Offloading in OPENMP
abstract
Rendering an image from a 3D scene requires a large amount of computation which grows exponentially with the complexity of the scene (e.g. number of objects and light sources). With the increasing demand of high definition content, 3D designers need to use high-performance computer systems to keep the rendering time acceptable. Since owning computer clusters is expensive, designers usually rent computing power directly from cloud service providers (e.g, AWS and Azure). However, even though many cloud providers already propose dedicated rendering services, integrating them within the standard workflow of modeling softwares can become a complex and cumbersome task. It typically requires exporting the project from the design software, dealing with various access control mechanisms from different clouds to upload the project, and executing the rendering remotely through command-line. Offloading computation to the cloud is a technique which can considerably simplify such tasks. To achieve that, this paper uses an extension of openMP 4.X to eliminate any major interactions with the end-user, while minimizing the complexity of cloud integration and optimizing the design workflow. It applies such approach to a ray-tracing application, a simplified version of the engines used by professional 3D modeling software (e.g. Blender). It automatically offloads the rendering process from the user computer to computer cluster within the Microsoft Azure cloud, brings the resulting images back after the computation ends and displays them directly on the screen of the user computer, thus providing a transparent programming model and good speed-ups over local execution.
Matheus Mortatti, Hervé Yviquel, Guido Araujo
SBAC-PAD2
2018 Cluster Programming using the OpenMP Accelerator Model
abstract
Computation offloading is a programming model in which program fragments (e.g., hot loops) are annotated so that their execution is performed in dedicated hardware or accelerator devices. Although offloading has been extensively used to move computation to GPUs, through directive-based annotation standards like OpenMP, offloading computation to very large computer clusters can become a complex and cumbersome task. It typically requires mixing programming models (e.g., OpenMP and MPI) and languages (e.g., C/C++ and Scala), dealing with various access control mechanisms from different cloud providers (e.g., AWS and Azure), and integrating all this into a single application. This article introduces computer cluster nodes as simple OpenMP offloading devices that can be used either from a local computer or from the cluster head-node. It proposes a methodology that transforms OpenMP directives to Spark runtime calls with fully integrated communication management, in a way that a cluster appears to the programmer as yet another accelerator device. Experiments using LLVM 3.8, OpenMP 4.5 on well known cloud infrastructures (Microsoft Azure and Amazon EC2) show the viability of the proposed approach, enable a thorough analysis of its performance, and make a comparison with an MPI implementation. The results show that although data transfers can impose overheads, cloud offloading from a local machine can still achieve promising speedups for larger granularity: up to 115× in 256 cores for the2MMbenchmark using 1GB sparse matrices. In addition, the parallel implementation of a complex and relevant scientific application reveals a 80× speedup on a 320 core machine when executed directly from the headnode of the cluster.
Hervé Yviquel, Lauro Cruz, Guido Araujo
ACM Trans. Archit. Code Optim.1
2017 The Cloud as an OpenMP Offloading Device
abstract
Computation offloading is a programming model in which program fragments (e.g. hot loops) are annotated so that their execution is performed in dedicated hardware or accelerator devices. Although offloading has been extensively used to move computation to GPUs, through directive-based annotation standards like OpenMP, offloading computation to very large computer clusters can become a complex and cumbersome task. It typically requires mixing programming models (e.g. OpenMP and MPI) and languages (e.g. C/C++ and Scala), dealing with various access control mechanisms from different clouds (e.g. AWS and Azure), and integrating all this into a single application. This paper introduces the cloud as a computation offloading device. It integrates OpenMP directives, cloud based map-reduce Spark nodes and remote communication management such that the cloud appears to the programmer as yet another device available in its local computer. Experiments using LLVM, OpenMP 4.5 and Amazon EC2 show the viability of the proposed approach and enable a thorough analysis of the performance and costs involved in cloud offloading. The results show that although data transfers can impose overheads, cloud offloading can still achieve promising speedups of up to 86x in 256 cores for the 2MM benchmark using 1GB matrices.
Hervé Yviquel, Guido Araujo
ICPP1
2014 Efficient software synthesis of dynamic dataflow programs
abstract
This paper introduces advanced software synthesis techniques that enhance the implementation of dynamic dataflow programs. These techniques have been implemented into open-source tools and demonstrated on well-known video decoders including one based on the new High Efficiency Video Coding (HEVC) standard. The results show an improvement of more than 100% of the frame-rate over previously proposed implementations, and achieve real-time decoding of high definition video sequences.
Hervé Yviquel, Alexandre Sanchez, Pekka Jääskeläinen, Jarmo Takala, Mickaël Raulet, Emmanuel Casseau
ICASSP1
2013 Orcc: multimedia development made easy
abstract
In this paper, we present Orcc, an open-source development environment that aims at enhancing multimedia development by offering all the advantages of dataflow programming: flexibility, portability and scalability. To do so, Orcc embeds two rich eclipse-based editors that provide an easy writing of dataflow applications, a simulator that allows quick validation of the written code, and a multi-target compiler that is able to translate any dataflow program, written in the RVC-CAL language, into an equivalent description in both hardware and software languages. Orcc has already been used to successfully write tens of multimedia applications, such as a video decoder supporting the new High Efficiency Video Coding standard, that clearly demonstrates the ability of the environment to develop complex applications. Moreover, results show scalable performances on multi-core platforms and achieve real-time decoding frame-rate on HD sequences.
Hervé Yviquel, Antoine Lorence, Khaled Jerbi, Gildas Cocherel, Alexandre Sanchez, Mickaël Raulet
ACM Multimedia1
2013 Automated design of networks of transport-triggered architecture processors using dynamic dataflow programs
Hervé Yviquel, Jani Boutellier, Mickaël Raulet, Emmanuel Casseau
Signal Process. Image Commun.1
2011 Just-in-time adaptive decoder engine: a universal video decoder based on MPEG RVC
abstract
In this paper, we introduce the Just-In-Time Adaptive Decoder Engine (Jade) project, which is shipped as part of the Open RVC-CAL Compiler (Orcc) project. Orcc provides a set of open-source software tools for managing decoders standardized within MPEG by the Reconfigurable Video Coding (RVC) experts. In this framework, Jade acts as a Virtual Machine for any decoder description that uses the MPEG RVC paradigm. Jade dynamically generates a native decoder representation suitable for X86, ARM and CELL platforms with a possibility of exploiting multi-core CPUs. Thus, according to the MPEG RVC decoder description coupled with a video coded stream, Jade can create, configure and re-configure video decompression algorithms adapting to the video content.
Jérôme Gorin, Hervé Yviquel, Françoise J. Prêteux, Mickaël Raulet
ACM Multimedia2