VLDB 2026 Research / reviewers in the wild / expert
Pablo de Oliveira Castro
dblp:37/8006
· DBLP profile ↗
20ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0001-9007-6145ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 4 first-author · 3 since 2021Theory of computation · 7 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Harnessing MPI mutations for AI error detectionabstractMPI errors are challenging to identify despite the significant number of expert verification tools. Dynamic tools (i.e., requiring profiling) are computationally expensive and accurate in error detection, whereas static analysis (i.e., operating at source code or compilation) is computationally cheap but less accurate. Interestingly, the recent success of AI and LLMs offers an alternative to increase static analysis accuracy while preserving its low overhead. Yet current methods remain difficult to benchmark, too general, and poorly adapted to the specific challenges of high-performance computing. Asia Auville, Tim Jammer, Eric Petit 0002, Pablo de Oliveira Castro, Emmanuelle Saillard, Mihail Popov |
ICS | 4 |
| 2026 | Enabling mixed-precision in spectral element codesabstractMixed-precision computing has the potential to significantly reduce the cost of exascale computations, but determining when and how to implement it in programs can be challenging. In this article, we propose a methodology for enabling mixed-precision with the help of computer arithmetic tools, roofline model, and computer arithmetic techniques. As case studies, we consider Nekbone (Nek5000 developers), a mini-application for the Computational Fluid Dynamics (CFD) solver Nek5000 (Fischer et al.), and a modern Neko (Jansson et al., 2024) CFD application. With the help of the Verificarlo (Denis et al., 2016) tool and computer arithmetic techniques, we introduce a strategy to address stagnation issues in the preconditioned Conjugate Gradient method in Nekbone and apply these insights to implement a mixed-precision version of Neko. We evaluate the derived mixed-precision versions of these codes by combining metrics in three dimensions: accuracy, time-to-solution, and energy-to-solution. Notably, mixed-precision in Nekbone reduces time-to-solution by roughly 1.62x and energy-to-solution by 2.43x on MareNostrum 5, while in the real-world Neko application, the gain is up to 1.3x in both time and energy, with the accuracy that matches double-precision results. Pablo de Oliveira Castro, Paolo Bientinesi, Niclas Jansson, Roman Iakymchuk |
Future Gener. Comput. Syst. | 2 |
| 2025 | Noise Injection for Performance Bottleneck Analysis
Aurélien Delval, Pablo de Oliveira Castro, William Jalby, Etienne Renault |
Euro-Par (1) | 2 |
| 2022 | The Positive Effects of Stochastic Rounding in Numerical AlgorithmsabstractRecently, stochastic rounding (SR) has been implemented in specialized hardware but most current computing nodes do not yet support this rounding mode. Several works empirically illustrate the benefit of stochastic rounding in various fields such as neural networks and ordinary differential equations. For some algorithms, such as summation, inner product or matrix-vector multiplication, it has been proved that SR provides probabilistic error bounds better than the traditional deterministic bounds. In this paper, we extend this theoretical ground for a wider adoption of SR in computer architecture. First, we analyze the biases of the two SR modes: SR-nearness and SR-up-or-down. We demonstrate on a case-study of Euler's forward method that IEEE-754 default rounding modes and SR-up-or-down accumulate rounding errors across iterations and that SR-nearness, being unbiased, does not. Second, we prove a $O(\sqrt{n})$ probabilistic bound on the forward error of Horner's polynomial evaluation method with SR, improving on the known deterministic O(n) bound. El-Mehdi El Arar, Devan Sohier, Pablo de Oliveira Castro, Eric Petit 0002 |
ARITH | 3 |
| 2021 | A Study of the Effects and Benefits of Custom-Precision Mathematical Libraries for HPC CodesabstractPublished in "IEEE Transactions on Emerging Topics in Computing, Volume: 9, Issue: 3, JulySeptember 2021" and orally presented at ARITH 2021. Emeric Brun, David Defour, Pablo de Oliveira Castro, Matei Istoan, Davide Mancusi, Eric Petit 0002, Alan Vaquet |
ARITH | 3 |
| 2021 | Shadow computation with BFloat16 to estimate the numerical accuracy of summationsabstractIn this article, we propose to exploit the new computational capability offered by the Bfloat16 representation format to perform shadow computations and compute estimations of the relative error. We demonstrate and evaluate the assumptions under which shadow computation is valid for the summation problem. David Defour, Pablo de Oliveira Castro, Matei Istoan, Eric Petit 0002 |
ARITH | 2 |
| 2021 | Confidence Intervals for Stochastic ArithmeticabstractQuantifying errors and losses due to the use of Floating-point (FP) calculations in industrial scientific computing codes is an important part of the Verification, Validation, and Uncertainty Quantification process. Stochastic Arithmetic is one way to model and estimate FP losses of accuracy, which scales well to large, industrial codes. It exists in different flavors, such as CESTAC or MCA, implemented in various tools such as CADNA, Verificarlo, or Verrou. These methodologies and tools are based on the idea that FP losses of accuracy can be modeled via randomness. Therefore, they share the same need to perform a statistical analysis of programs results to estimate the significance of the results. In this article, we propose a framework to perform a solid statistical analysis of Stochastic Arithmetic. This framework unifies all existing definitions of the number of significant digits (CESTAC and MCA), and also proposes a new quantity of interest: the number of digits contributing to the accuracy of the results. Sound confidence intervals are provided for all estimators, both in the case of normally distributed results, and in the general case. The use of this framework is demonstrated by two case studies of industrial codes: Europlexus and code_aster. Devan Sohier, Pablo de Oliveira Castro, François Févotte, Bruno Lathuilière, Eric Petit 0002, Olivier Jamond |
ACM Trans. Math. Softw. | 2 |
| 2020 | Custom-Precision Mathematical Library Explorations for Code Profiling and OptimizationabstractThe typical processors used for scientific computing have fixed-width data-paths. This implies that mathematical libraries were specifically developed to target each of these fixed precisions (binary16, binary32, binary64). However, to address the increasing energy consumption and throughput requirements of scientific applications, library and hardware designers are moving beyond this one-size-fits-all approach. In this article we propose to study the effects and benefits of using user-defined floating-point formats and target accuracies in calculations involving mathematical functions. Our tool collects input-data profiles and iteratively explores lower precisions for each call-site of a mathematical function in user applications. This profiling data will be a valuable asset for specializing and fine-tuning mathematical function implementations for a given application. We demonstrate the tool's capabilities on SGP4, a satellite tracking application. The profile data shows the potential for specialization and provides insight into answering where it is useful to provide variable-precision designs for elementary function evaluation. David Defour, Pablo de Oliveira Castro, Matei Istoan, Eric Petit 0002 |
ARITH | 2 |
| 2019 | Automatic Exploration of Reduced Floating-Point Representations in Iterative Methods
Yohan Chatelain, Eric Petit 0002, Pablo de Oliveira Castro, Ghislain Lartigue, David Defour |
Euro-Par | 3 |
| 2018 | VeriTracer: Context-enriched tracer for floating-point arithmetic analysisabstractVeriTracer automatically instruments a code and traces the accuracy of floating-point variables over time. VeriTracer enriches the visual traces with contextual information such as the call site path in which a value was modified. Contextual information is important to understand how the floating-point errors propagate in complex codes. VeriTracer is implemented as an LLVM compiler tool on top of Verificarlo. We demonstrate how VeriTracer can detect accuracy loss and quantify the impact of using a compensated algorithm on ABINIT, an industrial HPC application for Ab Initio quantum computation. Yohan Chatelain, Pablo de Oliveira Castro, Eric Petit 0002, David Defour, Jordan Bieder, Marc Torrent 0001 |
ARITH | 2 |
| 2017 | Piecewise holistic autotuning of parallel programs with CEREabstractSummary Current architecture complexity requires fine tuning of compiler and runtime parameters to achieve best performance. Autotuning substantially improves default parameters in many scenarios, but it is a costly process requiring long iterative evaluations. We propose an automatic piecewise autotuner based on CERE (Codelet Extractor and REplayer). CERE decomposes applications into small pieces called codelets: Each codelet maps to a loop or to an OpenMP parallel region and can be replayed as a standalone program. Codelet autotuning achieves better speedups at a lower tuning cost. By grouping codelet invocations with the same performance behavior, CERE reduces the number of loops or OpenMP regions to be evaluated. Moreover, unlike whole‐program tuning, CERE customizes the set of best parameters for each specific OpenMP region or loop. We demonstrate the CERE tuning of compiler optimizations, number of threads, thread affinity, and scheduling policy on both nonuniform memory access and heterogeneous architectures. Over the NAS benchmarks, we achieve an average speedup of 1.08× after tuning. Tuning a codelet is 13× cheaper than whole‐program evaluation and predicts the tuning impact with a 94.7% accuracy. Similarly, exploring thread configurations and scheduling policies for a Black‐Scholes solver on an heterogeneous big.LITTLE architecture is over 40× faster using CERE. Mihail Popov, Chadi Akel, Yohan Chatelain, William Jalby, Pablo de Oliveira Castro |
Concurr. Comput. Pract. Exp. | 5 |
| 2016 | Verificarlo: Checking Floating Point Accuracy through Monte Carlo ArithmeticabstractNumerical accuracy of floating point computation is a well studied topic which has not made its way to the end-user in scientific computing. Yet, it has become a critical issue with the recent requirements for code modernization to harness new highly parallel hardware and perform higher resolution computation. To democratize numerical accuracy analysis, it is important to propose tools and methodologies to study large use cases in a reliable and automatic way. In this paper, we propose verificarlo, an extension to the LLVM compiler to automatically use Monte Carlo Arithmetic in a transparent way for the end-user. It supports all the major languages including C, C++, and Fortran. Unlike source-to-source approaches, our implementation captures the influence of compiler optimizations on the numerical accuracy. We illustrate how Monte Carlo Arithmetic using the verificarlo tool outperforms the existing approaches on various use cases and is a step toward automatic numerical analysis. Christophe Denis, Pablo de Oliveira Castro, Eric Petit 0002 |
ARITH | 2 |
| 2016 | Piecewise Holistic Autotuning of Compiler and Runtime Parameters
Mihail Popov, Chadi Akel, William Jalby, Pablo de Oliveira Castro |
Euro-Par | 4 |
| 2015 | PCERE: Fine-Grained Parallel Benchmark Decomposition for Scalability PredictionabstractEvaluating the strong scalability of OpenMP applications is a costly and time-consuming process. It traditionally requires executing the whole application multiple times with different number of threads. We propose the Parallel Codelet Extractor and REplayer (PCERE), a tool to reduce the cost of scalability evaluation. PCERE decomposes applications into small pieces called codelets: each codelet maps to an OpenMP parallel region and can be replayed as a standalone program. To accelerate scalability prediction, PCERE replays codelets while varying the number of threads. Prediction speedup comes from two key ideas. First, the number of invocations during replay can be significantly reduced. Invocations that have the same performance are grouped together and a single representative is replayed. Second, sequential parts of the programs do not need to be replayed for each different thread configuration. PCERE codelets can be captured once and replayed accurately on multiple architectures, enabling cross-architecture parallel performance prediction. We evaluate PCERE on a C version of the NAS 3.0 Parallel Benchmarks (NPB). We achieve an average speed-up of 25 × on evaluating OpenMP applications scalability with an average error of 4.9% (median error of 1.7%). Mihail Popov, Chadi Akel, Florent Conti, William Jalby, Pablo de Oliveira Castro |
IPDPS | 5 |
| 2015 | CERE: LLVM-Based Codelet Extractor and REplayer for Piecewise Benchmarking and OptimizationabstractThis article presents Codelet Extractor and REplayer (CERE), an open-source framework for code isolation. CERE finds and extracts the hotspots of an application as isolated fragments of code, called codelets . Codelets can be modified, compiled, run, and measured independently from the original application. Code isolation reduces benchmarking cost and allows piecewise optimization of an application. Unlike previous approaches, CERE isolates codes at the compiler Intermediate Representation (IR) level. Therefore CERE is language agnostic and supports many input languages such as C, C++, Fortran, and D. CERE automatically detects codelets invocations that have the same performance behavior. Then, it selects a reduced set of representative codelets and invocations, much faster to replay, which still captures accurately the original application. In addition, CERE supports recompiling and retargeting the extracted codelets. Therefore, CERE can be used for cross-architecture performance prediction or piecewise code optimization. On the SPEC 2006 FP benchmarks, CERE codelets cover 90.9% and accurately replay 66.3% of the execution time. We use CERE codelets in a realistic study to evaluate three different architectures on the NAS benchmarks. CERE accurately estimates each architecture performance and is 7.3 × to 46.6 × cheaper than running the full benchmark. Pablo de Oliveira Castro, Chadi Akel, Eric Petit 0002, Mihail Popov, William Jalby |
ACM Trans. Archit. Code Optim. | 1 |
| 2014 | Fine-grained Benchmark Subsetting for System Selection
Pablo de Oliveira Castro, Yuriy Kashnikov, Chadi Akel, Mihail Popov, William Jalby |
CGO | 1 |
| 2013 | Is Source-Code Isolation Viable for Performance Characterization?abstractSource-code isolation finds and extracts the hotspots of an application as independent isolated fragments of code, called codelets. Codelets can be modified, compiled, run, and measured independently from the original application. Source-code isolation reduces benchmarking cost and allows piece-wise optimization of an application. Source-code isolation is faster than whole-program benchmarking and optimization since the user can concentrate only on the bottlenecks. This paper examines the viability of using isolated codelets in place of the original application for performance characterization and optimization. On the NAS benchmarks, we show that codelets capture 92.3% of the original execution time. We present a set of techniques for keeping codelets as faithful as possible to the original hotspots: 63.6% of the codelets have the same assembly as the original hotspots and 81.6% of the codelets have the same run time performance as the original hotspots. Chadi Akel, Yuriy Kashnikov, Pablo de Oliveira Castro, William Jalby |
ICPP | 3 |
| 2013 | Adaptive sampling for performance characterization of application kernelsabstractSUMMARY Characterizing performance is essential to optimize programs and architectures. The open source Adaptive Sampling Kit (ASK) measures the performance trade‐off in large design spaces. Exhaustively sampling all sets of parameters is computationally intractable. Therefore, ASK concentrates exploration in the most irregular regions of the design space through multiple adaptive sampling strategies. The paper presents the ASK architecture and a set of adaptive sampling strategies, including a new approach called Hierarchical Variance Sampling. ASK's usage is demonstrated on three performance characterization problems: memory stride accesses, Jacobian stencil code, and an industrial seismic application using 3D stencils. ASK builds accurate models of performance with a small number of measures. It considerably reduces the cost of performance exploration. For instance, the Jacobian stencil code design space, which has more than 31 × 108 combinations of parameters, is accurately predicted using only 1500 combinations. Copyright © 2013 John Wiley & Sons, Ltd. Pablo de Oliveira Castro, Eric Petit 0002, Asma Farjallah, William Jalby |
Concurr. Comput. Pract. Exp. | 1 |
| 2012 | ASK: Adaptive Sampling Kit for Performance Characterization
Pablo de Oliveira Castro, Eric Petit 0002, Jean Christophe Beyler, William Jalby |
Euro-Par | 1 |
| 2010 | A Multidimensional Array Slicing DSL for Stream ProgrammingabstractStream languages offer a simple multi-core programming model and achieve good performance. Yet expressing data rearrangement patterns (like a matrix block decomposition) in these languages is verbose and error prone. In this paper, we propose a high-level programming language to elegantly describe n-dimensional data reorganization patterns. We show how to compile it to stream languages. Pablo de Oliveira Castro, Stéphane Louise, Denis Barthou |
CISIS | 1 |