VLDB 2026 Research / reviewers in the wild / expert
Mihail Popov
dblp:141/4409
· DBLP profile ↗
16ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0002-3498-8147ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 4 first-author · 7 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Harnessing MPI mutations for AI error detectionabstractMPI errors are challenging to identify despite the significant number of expert verification tools. Dynamic tools (i.e., requiring profiling) are computationally expensive and accurate in error detection, whereas static analysis (i.e., operating at source code or compilation) is computationally cheap but less accurate. Interestingly, the recent success of AI and LLMs offers an alternative to increase static analysis accuracy while preserving its low overhead. Yet current methods remain difficult to benchmark, too general, and poorly adapted to the specific challenges of high-performance computing. Asia Auville, Tim Jammer, Eric Petit 0002, Pablo de Oliveira Castro, Emmanuelle Saillard, Mihail Popov |
ICS | 6 |
| 2026 | TOTO: Transparent I/O Tuning for HPC ApplicationsabstractHigh-performance computing applications rely on parallel file systems, where I/O performance is strongly affected by configuration parameters such as stripe count. However, the ideal stripe count is highly application- and system-dependent, making it difficult to predict and rarely tuned in practice. As a result, substantial I/O performance potential remains unexplored. We present TOTO, a transparent tool that improves I/O performance without requiring application modifications. TOTO intercepts POSIX calls, characterizes application behavior, and uses a machine learning model to select an appropriate stripe count, even for already opened files. We also introduce an allocation algorithm that balances performance and resource occupation, and describe a methodology for training the model once per system using limited data. Our results show that TOTO can improve I/O performance by up to 4.6 × compared to using a default stripe count, while imposing an overhead of at most \(8\%\). Moreover, compared to the state of the art, TOTO can optimize more applications with a lower resource occupation, which is expected to decrease contention in the I/O infrastructure. Francieli Zanon Boito, Luan Teylo, Mihail Popov, Laora Aimi, Alexis Bandet, Laércio Lima Pilla, Guillaume Pallez |
ICS | 3 |
| 2025 | A Deep Look into the Temporal I/O Behavior of HPC ApplicationsabstractThe increasing gap between compute and I/O speeds in high-performance computing (HPC) systems imposes the need for techniques to improve applications' I/O performance. Such techniques must rely on assumptions about I/O behavior in order to efficiently allocate I/O resources such as burst buffers, to schedule accesses to the shared parallel file system or to delay certain applications at the batch scheduler level to prevent contention, for instance. In this paper, we verify these common assumptions about I/O behavior, specifically about temporal behavior, using over 440,000 traces from real HPC systems. By combining traces from diverse systems, we characterize the behaviors observed in real HPC workloads. Among other findings, we show that I/O activity tends to last for a few seconds, and that periodic jobs are the minority, but responsible for a large portion of the I/O time. Furthermore, we make projections for the expected improvement yielded by popular approaches for I/O performance improvement. Our work provides valuable insights to everyone working to alleviate the I/O bottleneck in HPC. Francieli Zanon Boito, Luan Teylo, Mihail Popov, Théo Jolivel, Francois Tessier, Jakob Lüttgau, Julien Monniot, Ahmad Tarraf, Andre Ramos Carneiro, Carla Osthoff |
IPDPS | 3 |
| 2025 | Compiler, Runtime, and Hardware Parameters Design Space ExplorationabstractHPC systems are increasingly complex with many tunable parameters impacting applications' metrics-e.g., performance, energy consumption. The main challenges of these systems are finding the appropriate configuration per application on any given system and understanding how the configurations affect applications' metrics on a system. Both can be addressed with design space exploration (DSE). However, exploring all the configurations available is costly due to the long execution and setup times of these executions. Indeed, it requires instrumenting the applications to collect data, compiling them with different options and setting the parameters for each execution. DSE algorithms can greatly reduce the exploration time by guiding which configuration to execute next to reach the objective without evaluating all the configurations. A DSE study thus requires implementing an exploration algorithm and automating parameters setting, application instrumentation and compilation, and metrics collection. This represents a huge overhead to the actual study, yet most DSE studies still do it from scratch. To alleviate the setup cost, we propose a unified methodology to perform the exploration and implement it in the CORHPEX framework to setup configurations with compiler, runtime, and hardware parameters, efficiently and flexibly. The framework enables choosing the exploration strategy, the design space to study, the applications to execute and the metrics to collect independently while involving little coding overhead. It is extensible with custom exploration algorithms and data readers. We demonstrate the versatility and robustness of our framework on parallel codes, including NAS, Rodinia, LULESH benchmarks, and real-world applications, on two systems exposing different parameters with various DSE techniques and goals. We show that working with CORHPEX enables getting insights on code optimization strategies by using exploration algorithms that can speedup the execution by a factor of 10X while preserving 95% the possible gains. Finally, we demonstrate the framework's potential for more advanced studies by training surrogate models of complex HPC applications achieving over 93% accuracy. Lana Scravaglieri, Ani Anciaux-Sedrakian, Olivier Aumage, Thomas Guignon, Mihail Popov |
IPDPS | 5 |
| 2024 | MPI Errors Detection using GNN Embedding and Vector Embedding over LLVM IRabstractIdentifying errors in parallel MPI programs is a challenging task. Despite the growing number of verification tools, debugging parallel programs remains a significant challenge. This paper is the first to utilize embedding and deep learning graph neural networks (GNNs) to tackle the issue of identifying bugs in MPI programs. Specifically, we have designed and developed two models that can determine, from a code’s LLVM Intermediate Representation (IR), whether the code is correct or contains a known MPI error.We tested our models using two dedicated MPI benchmark suites for verification: MBI and MPI-CorrBench. By training and validating our models on the same benchmark suite, we achieved a prediction accuracy of 92% in detecting error types. Additionally, we trained and evaluated our models on distinct benchmark suites (e.g., transitioning from MBI to MPI-CorrBench) and achieved a promising accuracy of over 80%. Finally, we investigated the interaction between different MPI errors and quantified our models generalization capabilities over new unseen errors. This involved removing errors types during training and assessing whether our models could still predict them. The detection accuracy of removed errors vary significantly between 20% to 80%, indicating connected error patterns. Jad El Karchi, Hanze Chen, Ali TehraniJamsaz, Ali Jannesari, Mihail Popov, Emmanuelle Saillard |
IPDPS | 5 |
| 2023 | Optimizing performance and energy across problem sizes through a search space exploration and machine learning
Lana Scravaglieri, Mihail Popov, Laércio Lima Pilla, Amina Guermouche, Olivier Aumage, Emmanuelle Saillard |
J. Parallel Distributed Comput. | 2 |
| 2022 | Learning Intermediate Representations using Graph Neural Networks for NUMA and Prefetchers OptimizationabstractThere is a large space of NUMA and hardware prefetcher configurations that can significantly impact the performance of an application. Previous studies have demonstrated how a model can automatically select configurations based on the dynamic properties of the code to achieve speedups. This paper demonstrates how the static Intermediate Representation (IR) of the code can guide NUMA/prefetcher optimizations without the prohibitive cost of performance profiling. We propose a method to create a comprehensive dataset that includes a diverse set of intermediate representations along with optimum configurations. We then apply a graph neural network model in order to validate this dataset. We show that our static intermediate representation based model achieves 80 % of the performance gains provided by expensive dynamic performance profiling based strategies. We further develop a hybrid model that uses both static and dynamic information. Our hybrid model achieves the same gains as the dynamic models but at a reduced cost by only profiling 30 % of the programs. Ali TehraniJamsaz, Mihail Popov, Akash Dutta, Emmanuelle Saillard, Ali Jannesari |
IPDPS | 2 |
| 2022 | Analysing and Predicting Energy Consumption of Garbage Collectors in OpenJDKabstractSustainable computing needs energy-efficient software. This paper explores the potential of leveraging the nature of software written in managed languages: increasing energy efficiency by changing a program’s memory management strategy without altering a single line of code. To this end, we perform comprehensive energy profiling of 35 Java applications across four benchmarks. In many cases, we find that it is possible to save energy by replacing the default G1 collector with another without sacrificing performance. Furthermore, potential energy savings can be even higher if performance regressions are permitted. Inspired by these results, we study what the most energy-efficient GCs are to help developers prune the search space for energy profiling at a low cost. Finally, we show that machine learning can be successfully applied to the problem of finding an energy-efficient GC configuration for an application, reducing the cost even further. Marina Shimchenko, Mihail Popov, Tobias Wrigstad |
MPLR | 2 |
| 2020 | Modeling and optimizing NUMA effects and prefetching with machine learningabstractBoth NUMA thread/data placement and hardware prefetcher configuration have significant impacts on HPC performance. Optimizing both together leads to a large and complex design space that has previously been impractical to explore at runtime. Isaac Sánchez Barrera, David Black-Schaffer, Marc Casas, Miquel Moretó, Anastasiia Stupnikova, Mihail Popov |
ICS | 6 |
| 2019 | Efficient thread/page/parallelism autotuning for NUMA systemsabstractCurrent multi-socket systems have complex memory hierarchies with significant Non-Uniform Memory Access (NUMA) effects: memory performance depends on the location of the data and the thread. This complexity means that thread- and data-mappings have a significant impact on performance. However, it is hard to find efficient data mappings and thread configurations due to the complex interactions between applications and systems. Mihail Popov, Alexandra Jimborean, David Black-Schaffer |
ICS | 1 |
| 2018 | The Long and Winding Road Toward Efficient High-Performance ComputingabstractThe major challenge to Exaflop computing, and more generally, efficient high-end computing, is in finding the best “matches” between advanced hardware capabilities and the software used to program applications, so that top performance will be achieved. Several benchmarks show very disappointing performance progress over the last decade, clearly indicating a mismatch between hardware and software. To remedy this problem, it is important that key performance enablers at the software level-autotuning, performance analysis tools, full application optimization-are understood. For each area, we highlight major limitations and most promising approaches to reaching better performance and energy levels. Finally, we conclude by analyzing hardware and software design, trying to pave the way for more tightly integrated hardware and software codesign. William Jalby, David J. Kuck, Allen D. Malony, Michel Masella, Abdelhafid Mazouz, Mihail Popov |
Proc. IEEE | 6 |
| 2017 | Piecewise holistic autotuning of parallel programs with CEREabstractSummary Current architecture complexity requires fine tuning of compiler and runtime parameters to achieve best performance. Autotuning substantially improves default parameters in many scenarios, but it is a costly process requiring long iterative evaluations. We propose an automatic piecewise autotuner based on CERE (Codelet Extractor and REplayer). CERE decomposes applications into small pieces called codelets: Each codelet maps to a loop or to an OpenMP parallel region and can be replayed as a standalone program. Codelet autotuning achieves better speedups at a lower tuning cost. By grouping codelet invocations with the same performance behavior, CERE reduces the number of loops or OpenMP regions to be evaluated. Moreover, unlike whole‐program tuning, CERE customizes the set of best parameters for each specific OpenMP region or loop. We demonstrate the CERE tuning of compiler optimizations, number of threads, thread affinity, and scheduling policy on both nonuniform memory access and heterogeneous architectures. Over the NAS benchmarks, we achieve an average speedup of 1.08× after tuning. Tuning a codelet is 13× cheaper than whole‐program evaluation and predicts the tuning impact with a 94.7% accuracy. Similarly, exploring thread configurations and scheduling policies for a Black‐Scholes solver on an heterogeneous big.LITTLE architecture is over 40× faster using CERE. Mihail Popov, Chadi Akel, Yohan Chatelain, William Jalby, Pablo de Oliveira Castro |
Concurr. Comput. Pract. Exp. | 1 |
| 2016 | Piecewise Holistic Autotuning of Compiler and Runtime Parameters
Mihail Popov, Chadi Akel, William Jalby, Pablo de Oliveira Castro |
Euro-Par | 1 |
| 2015 | PCERE: Fine-Grained Parallel Benchmark Decomposition for Scalability PredictionabstractEvaluating the strong scalability of OpenMP applications is a costly and time-consuming process. It traditionally requires executing the whole application multiple times with different number of threads. We propose the Parallel Codelet Extractor and REplayer (PCERE), a tool to reduce the cost of scalability evaluation. PCERE decomposes applications into small pieces called codelets: each codelet maps to an OpenMP parallel region and can be replayed as a standalone program. To accelerate scalability prediction, PCERE replays codelets while varying the number of threads. Prediction speedup comes from two key ideas. First, the number of invocations during replay can be significantly reduced. Invocations that have the same performance are grouped together and a single representative is replayed. Second, sequential parts of the programs do not need to be replayed for each different thread configuration. PCERE codelets can be captured once and replayed accurately on multiple architectures, enabling cross-architecture parallel performance prediction. We evaluate PCERE on a C version of the NAS 3.0 Parallel Benchmarks (NPB). We achieve an average speed-up of 25 × on evaluating OpenMP applications scalability with an average error of 4.9% (median error of 1.7%). Mihail Popov, Chadi Akel, Florent Conti, William Jalby, Pablo de Oliveira Castro |
IPDPS | 1 |
| 2015 | CERE: LLVM-Based Codelet Extractor and REplayer for Piecewise Benchmarking and OptimizationabstractThis article presents Codelet Extractor and REplayer (CERE), an open-source framework for code isolation. CERE finds and extracts the hotspots of an application as isolated fragments of code, called codelets . Codelets can be modified, compiled, run, and measured independently from the original application. Code isolation reduces benchmarking cost and allows piecewise optimization of an application. Unlike previous approaches, CERE isolates codes at the compiler Intermediate Representation (IR) level. Therefore CERE is language agnostic and supports many input languages such as C, C++, Fortran, and D. CERE automatically detects codelets invocations that have the same performance behavior. Then, it selects a reduced set of representative codelets and invocations, much faster to replay, which still captures accurately the original application. In addition, CERE supports recompiling and retargeting the extracted codelets. Therefore, CERE can be used for cross-architecture performance prediction or piecewise code optimization. On the SPEC 2006 FP benchmarks, CERE codelets cover 90.9% and accurately replay 66.3% of the execution time. We use CERE codelets in a realistic study to evaluate three different architectures on the NAS benchmarks. CERE accurately estimates each architecture performance and is 7.3 × to 46.6 × cheaper than running the full benchmark. Pablo de Oliveira Castro, Chadi Akel, Eric Petit 0002, Mihail Popov, William Jalby |
ACM Trans. Archit. Code Optim. | 4 |
| 2014 | Fine-grained Benchmark Subsetting for System Selection
Pablo de Oliveira Castro, Yuriy Kashnikov, Chadi Akel, Mihail Popov, William Jalby |
CGO | 4 |