Dolores Miao

dblp:347/0128 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0002-7511-0269ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
GPUs and heterogeneous computing · 44% Performance modeling and evaluation · 44% High-performance computing · 13%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
GPU programming
0.912025
FloatGuard: Efficient Whole-Program Detection of Floating-Point Exceptions in AMD GPUs · HPDC 2025
Performance modeling and evaluation
program instrumentation
0.912025
FloatGuard: Efficient Whole-Program Detection of Floating-Point Exceptions in AMD GPUs · HPDC 2025
High-performance computing › scientific computing
scientific computing application
0.312025
FloatGuard: Efficient Whole-Program Detection of Floating-Point Exceptions in AMD GPUs · HPDC 2025

Methods — techniques the papers use, named apart from their topics

source-level instrumentation · 0.9debugger-guided execution · 0.9assembly-level instrumentation · 0.9
YearPublicationVenuePosition
2025 FloatGuard: Efficient Whole-Program Detection of Floating-Point Exceptions in AMD GPUs
abstract
Porting scientific applications across different GPU architectures introduces floating-point arithmetic variations that can affect reproducibility, making efficient detection and mitigation of exceptions like NaNs and infinities crucial. While NVIDIA has dominated the GPU market, AMD GPUs are increasingly used in HPC systems as well, yet existing floating-point exception detection frameworks focus on NVIDIA, leaving a gap for AMD GPUs. We present FloatGuard, the first framework for efficiently detecting floating-point exceptions in HIP programs running on AMD GPUs. FloatGuard leverages AMD GPU hardware registers to detect floating-point exceptions, overcoming the limitations of AMD's built-in trapping mechanisms through a novel algorithm that combines assembly- and source-level instrumentation with debugger-guided execution. We evaluate FloatGuard on 565 HIP programs, detecting floatingpoint exceptions in 507 cases with a slowdown ratio that increases at most linearly with the number of exceptions discovered. Furthermore, we analyze the impact of compiler optimizations on exceptions trapped, and compare FloatGuard with the state-of-the-art tool for detecting floating-point exceptions in CUDA programs, which further reveals key differences between AMD and NVIDIA's floating-point exception behaviors.
Dolores Miao, Ignacio Laguna, Cindy Rubio-González
HPDC1
2024 Input Range Generation for Compiler-Induced Numerical Inconsistencies
abstract
Compiler-induced numerical inconsistencies present a significant challenge when testing and verifying numerical software—they can arise in a variety of situations, such as when porting code to a new platform or when using a different compiler or optimization flag. While existing tools can identify the source code location that induce an inconsistency for a specific input, no techniques are available to find input ranges where inputs that trigger these inconsistencies exist. In this paper, we propose a multi-phase approach to detect unknown input ranges that induce such inconsistencies; we call them inconsistency-inducing inputs. Our approach combines input-partitioned and coverage-based input sampling, input clustering, and optimization algorithms. We implement our approach in the tool CIGEN, which finds inputs that trigger high compiler-induced inconsistencies in numerical programs and outputs a list of input ranges containing such inconsistency-inducing inputs. Our experimental evaluation show 53.4% improvement over the state of the art in finding inputs that trigger compiler-induced inconsistencies in 175 GNU Scientific Library (GSL) functions. We further examine a subset of the inconsistencies and discuss their characteristics and possible root causes.
Dolores Miao, Ignacio Laguna, Cindy Rubio-González
ICS1
2024 An automated OpenMP mutation testing framework for performance optimization
abstract
Performance optimization continues to be a challenge in modern HPC software. Existing performance optimization techniques, including profiling-based and auto-tuning techniques, fail to indicate program modifications at the source level thus preventing their portability across compilers. This paper describes Muppet, a new approach that identifies program modifications called mutations aimed at improving program performance. Muppet’s mutations help developers reason about performance defects and missed opportunities to improve performance at the source code level. In contrast to compiler techniques that optimize code at intermediate representations (IR), Muppet uses the idea of source-level mutation testing to relax correctness constraints and automatically discover optimization opportunities that otherwise are not feasible using the IR. We demonstrate the Muppet’s concept in the OpenMP programming model. Muppet generates a list of OpenMP mutations that alter the program parallelism in various ways, and is capable of running a variety of optimization algorithms such as delta debugging, Bayesian Optimization and decision tree optimization to find a subset of mutations which, when applied to the original program, cause the most speedup while maintaining program correctness. When Muppet is evaluated against a diverse set of benchmark programs and proxy applications, it is capable of finding sets of mutations that induce speedup in 75.9% of the evaluated programs.
Dolores Miao, Ignacio Laguna, Giorgis Georgakoudis, Konstantinos Parasyris, Cindy Rubio-González
Parallel Comput.1