Philippe Couvée

dblp:266/2231 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0002-8989-9408ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 I/O patterns modeling of HPC applications with call stacks for predictive prefetch
abstract
Modern high-performance computing (HPC) storage systems use heterogeneous storage technologies organized in tiers to find a compromise between capacity, performance, and cost. In these systems, prefetching is a common technique used to move the right data at the right moment from a slow to a fast tier to improve overall performance while using the costly high-performance tier only when needed. Effective prefetching requires precise knowledge of the application I/O patterns. This knowledge can be extracted through the source code, I/O tracing tools or I/O functions call stacks. State-of-the-art solutions based on the latter approach mainly focus on applications with regular I/O profiles to avoid scalability issues due to the grammar-based techniques used. In this paper, we present an approach based on I/O call stacks that models POSIX and STDIO I/O patterns for both regular and irregular applications, thanks to the use of directed graphs. We present different models usable for prefetching. Our models were used to predict the next I/O call stack on five real HPC applications with a prediction accuracy of up to 98%. Compared to the state-of-the-art Omnisc’IO, they incurred up to 120x lower model overhead (334 ns vs. 45 μ s on LAMMPS) and had a model size 10x to 15x smaller (463 B vs. 7 kB on LQCD).
Louis-Marie Nicolas, Salim Mimouni, Philippe Couvée, Jalil Boukhobza
Future Gener. Comput. Syst.3
2025 Inter-File Similarity for Online Data Placement Prediction in Hierarchical Storage
abstract
The growing disparity between computing speed and data access times poses a significant challenge to the management of data storage in large-scale supercomputers. To address these challenges, storage systems have evolved into hierarchical architectures with multiple levels that accommodate various hardware technologies. Each level offers a unique combination of performance, cost, and capacity. In this work, we introduce a paradigm shift in data management across the storage tiers by transitioning from block-level to file-level granularity to optimize data placement. Our online prediction model uses file characterization (input, log, checkpoint, work, and output) to predict file reuse with 97% accuracy, categorizing each file during usage and enabling proactive data management strategies. Using three real applications and a benchmark, we validate our work and demonstrate significant improvements in data hit rates, reaching 55% and 57% for LQCD and the benchmark applications, respectively. Even with lower performance improvements, our approach remains competitive with NEMO and NAMD, achieving a 3% gain over both LRU and LFU on NEMO and matching LRU performance while being 1% lower than LFU on NAMD.
Adrian Khelili, Sophie Robert-Hayek, Soraya Zertal, Philippe Couvée
HPCC4
2025 Investigating the Use of File Advisory Hints on Lustre and GPFS
abstract
International audience
Louis-Marie Nicolas, Salim Mimouni, Philippe Couvée, Jalil Boukhobza
SSDBM3
2021 A comparative study of black-box optimization heuristics for online tuning of high performance computing I/O accelerators
abstract
Summary High performance computing (HPC) applications' behaviors rely on highly configurable software environments and hardware devices. Finding their optimal parametrization is a complex task, as the size of their parametric space and the non‐linear behavior of HPC systems make hand‐tuning, theoretical modeling or exhaustive sampling unsuitable in most cases. In this article, we propose an online auto‐tuner that relies on black‐box optimization to find the optimal parametrization of input/output (I/O) accelerators for a given HPC application in a limited number of iterations, without making any assumption on the behavior of the tuned system. As many heuristics are available in the literature, we need to guarantee the quality of the tuning by selecting the most appropriate one. To do so, we provide a comparative study of the efficiency of three heuristics applied to tuning two I/O accelerators developed by the Atos company: a pure software accelerator (small read optimizer) and a mixed hardware‐software one (smart burst buffer). To select the most efficient heuristic for our use case, we define several new metrics to evaluate the quality of an auto‐tuner, in an online and offline settings. We find that genetic algorithms provide a faster convergence rate and a faster computation time but surrogate models provide a better score in terms of both distance to the optimum and trajectory stability. Overall, the obtained results show that auto‐tuning heuristics improve the execution time of applications used conjointly with both SRO and SBB accelerators.
Sophie Robert-Hayek, Soraya Zertal, Gregory Vaumourin, Philippe Couvée
Concurr. Comput. Pract. Exp.4
2019 uMMAP-IO: User-Level Memory-Mapped I/O for HPC
abstract
The integration of local storage technologies alongside traditional parallel file systems on HPC clusters, is expected to rise the programming complexity on scientific applications aiming to take advantage of the increased-level of heterogeneity. In this work, we present uMMAP-IO, a user-level memory-mapped I/O implementation that simplifies data management on multi-tier storage subsystems. Compared to the memory-mapped I/O mechanism of the OS, our approach features per-allocation configurable settings (e.g., segment size) and transparently enables access to a diverse range of memory and storage technologies, such as the burst buffer I/O accelerators. Preliminary results indicate that uMMAP-IO provides at least 5-10x better performance on representative workloads in comparison with the standard memory-mapped I/O of the OS, and approximately 20-50% degradation on average compared to using conventional memory allocations without storage support up to 8192 processes.
Sergio Rivas-Gomez, Alessandro Fanfarillo, Sébastien Valat, Christophe Laferriere, Philippe Couvée, Sai Narasimhamurthy, Stefano Markidis
HiPC5