José Monsalve Diaz

dblp:224/0615 · also Jose Manuel Monsalve Diaz, Jose Monsalve Diaz · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0001-6875-1685ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 4 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Profiler-Guided Execution of Recurrent OpenMP Task Graphs on Heterogeneous Clusters
abstract
Distributed task-based execution models are well-suited for parallelizing irregular applications across clusters. OpenMP Cluster (OMPC) extends the traditional OpenMP tasking model to support distributed memory systems, leveraging a HEFT-based scheduler to improve resource utilization. However, the efficiency of such a scheduler depends heavily on accurate estimates of task execution and communication costs – information that is often difficult to obtain reliably and efficiently. To address this limitation, we propose a novel scheduling framework that combines the recent taskgraph directive introduced in OpenMP 6.0 with partial online profiling of iterative applications. Our approach performs quasi-static scheduling by recording task graphs at runtime and selectively profiling representative iterations to estimate performance. This information is interpolated and fed back into the scheduler to enhance decision-making. We demonstrate that our framework can improve scheduling quality with minimal overhead, making it suitable for long-running or repetitive workloads commonly found in High-Performance Computing (HPC) applications. We achieve up to 20% speedup for the total application and 4× speedup for scheduling.
Rémy Neveu, Rodrigo Ceccato, Adrian Munera, Sara Royuela, José Monsalve Diaz, Hervé Yviquel
SBAC-PAD5
2023 Implementation of Dataflow Software Pipelining for Codelet Model
Siddhisanket Raskar, José Monsalve Diaz, Thomas Applencourt, Kalyan Kumaran, Guang R. Gao
ICPE2
2022 Efficient Execution of OpenMP on GPUs
abstract
OpenMP is the preferred choice for CPU parallelism in High-Performance-Computing (HPC) applications written in C, C++, or Fortran. As HPC systems became heterogeneous, OpenMP introduced support for accelerator offloading via the target directive. This allowed porting existing (CPU) code onto GPUs, including well established CPU parallelism paradigms. However, there are architectural differences between CPU and GPU execution which make common patterns, like forking and joining threads, single threaded execution, or sharing of local (stack) variables, in general costly on the latter. So far it was left to the user to identify and avoid non-efficient code patterns, most commonly by writing their OpenMP offloading codes in a kernel-language style which resembles CUDA more than it does traditional OpenMP.In this work we present OpenMP-aware program analyses and optimizations that allow efficient execution of the generic, CPU-centric parallelism model provided by OpenMP on GPUs. Our implementation in LLVM/Clang maps various common OpenMP patterns found in real world applications efficiently to the GPU. As static analysis is inherently limited we provide actionable and informative feedback to the user about the performed and missed optimizations, together with ways for the user to annotate the program for better results. Our extensive evaluation using several HPC proxy applications shows significantly improved GPU kernel times and reduction in resources requirements, such as GPU registers.
Joseph Huber, Melanie Cornelius, Giorgis Georgakoudis, Shilei Tian, José Monsalve Diaz, Kuter Dinel, Barbara M. Chapman, Johannes Doerfert
CGO5
2022 Co-Designing an OpenMP GPU Runtime and Optimizations for Near-Zero Overhead Execution
abstract
GPU accelerators are ubiquitous in modern HPC systems. To program them, users have the choice between vendor-specific, native programming models, such as CUDA, which provide simple parallelism semantics with minimal runtime support, or portable alternatives, such as OpenMP, which offer rich parallel semantics and feature an extensive runtime library to support execution. While the operations of such a runtime can easily limit performance and drain resources, it was to some degree regarded an unavoidable overhead. In this work we present a co-design methodology for optimizing applications using a specifically crafted OpenMP GPU runtime such that most use cases induce near-zero overhead. Specifically, our approach exposes runtime semantics and state to the compiler such that optimization effectively eliminating abstractions and runtime state from the final binary. With the help of user provided assumptions we can further optimize common patterns that otherwise increase resource consumption. We evaluated our prototype build on top of the LLVM/OpenMP GPU offloading infrastructure with multiple HPC proxy applications and benchmarks. Comparison of CUDA, the original OpenMP runtime, and our co-designed alternative show that, by our approach, performance is significantly improved and resource consumption is significantly lowered. Oftentimes we can closely match the CUDA implementation without sacrificing the versatility and portability of OpenMP.
Johannes Doerfert, Atmn Patel, Joseph Huber, Shilei Tian, José Monsalve Diaz, Barbara M. Chapman, Giorgis Georgakoudis
IPDPS5
2021 swFLOW: A large-scale distributed framework for deep learning on Sunway TaihuLight supercomputer
Mingfan Li, Junshi Chen 0003, José Monsalve Diaz, Rongfen Lin, Guang R. Gao, Hong An
Inf. Sci.4
2019 Toward A High-Performance Emulation Platformfor Brain-Inspired Intelligent SystemsExploring Dataflow-Based Execution Model and Beyond
abstract
Brain-inspired computing is a novel computing technology based on neural morphological engineering, which draws lessons from methods of human brain information processing and storage. Combining with the high-performance computing (HPC) platform, they constitute the foundation of general artificial intelligence. However, current brain HPC platforms generally suffer from slow speed, poor scalability, and high energy consumption, which severely restrain its potential and circumscribe the development of general artificial intelligence. The dataflow model was first proposed in the 1970s, providing a novel idea for the development of HPC. In addition, the dataflow model shares similar information processing mechanisms with human's neural system, which makes dataflow models naturally suit the emulation of brain-inspired computing. Based on the contemporary progress of the dataflow model, the Codelet model was proposed. Through a fine-grained asynchronous program execution and resource allocation, the Codelet model successfully realized the distributed computing on the heterogeneous system, effectively improved the computing power and speed, and open up a new path to overcome the shortcomings of the existing high-performance computing technology. We propose a dataflow-based emulation platform, aiming at providing high-performance computing technology support for general brain-inspired intelligent system, as well as using characteristics of dataflow models to fully explore the potential of brain-inspired intelligence. As an example, we will select a convolutional neural network (LeNet5) that already has a spectacular user base to initially verify the superiority and feasibility of our proposal.
Sihan Zeng, José Monsalve Diaz, Siddhisanket Raskar
COMPSAC (2)2
2019 Analysis of OpenMP 4.5 Offloading in Implementations: Correctness and Overhead
José Monsalve Diaz, Kyle Friedline, Swaroop Pophale, Oscar R. Hernandez, David E. Bernholdt, Sunita Chandrasekaran
Parallel Comput.1
2015 Dynamic CPU Resource Allocation in Containerized Cloud Environments
abstract
In recent years, lighter-weight virtualization solutions have begun to emerge as an alternative to virtual machines. Because these solutions are still in their infancy, however, several research questions remain open in terms of how to effectively manage computing resources. One important problem is the management of resources in the event of overutilization. For some applications, overutilization can severely affect performance. We provide a solution to this problem by extending the concept of timeslicing to the level of virtualization container. Through this approach we can control and mitigate some of the more detrimental performance effects oversubscription. Our results show significant improvement over standard scheduling with Docker.
José Monsalve Diaz, Aaron Myles Landwehr, Michela Taufer
CLUSTER1
2015 Improving MPSoC reliability through adapting runtime task schedule based on time-correlated fault behavior
Laura A. Rozo Duque, José Monsalve Diaz, Chengmo Yang
DATE2