Guillermo Indalecio Fernández

dblp:117/7605 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
0since 2021 · last 2019
0000-0001-7727-1704ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
High-performance computing · 23% Parallel and multicore computing · 23% Emerging computing paradigms · 23%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › memory access optimization
data movement reduction
0.412019
Optimizing the data movement in quantum transport simulations via data-centric parallel programming · SC 2019
High-performance computing
multi-physics simulation
0.412019
A data-centric approach to extreme-scale ab initio dissipative quantum transport simulations · SC 2019
Emerging computing paradigms › quantum computing › quantum simulation
quantum transport simulation
0.412019
Optimizing the data movement in quantum transport simulations via data-centric parallel programming · SC 2019
Parallel and multicore computing › parallel programming models
task-based programming
0.412019
A data-centric approach to extreme-scale ab initio dissipative quantum transport simulations · SC 2019
Electronic design automation
electrothermal simulation
0.112019
Optimizing the data movement in quantum transport simulations via data-centric parallel programming · SC 2019

Methods — techniques the papers use, named apart from their topics

regent programming language · 0.8legion programming system · 0.8data layout transformation · 0.4communication avoidance · 0.4
YearPublicationVenuePosition
2019 A data-centric approach to extreme-scale ab initio dissipative quantum transport simulations
abstract
The Predictive Science Academic Alliance Program (PSAAP) II at Stanford University is developing an exascale-ready multi-physics solver to investigate particle-laden turbulent flows in a radiation environment for solar energy receiver applications. In order to simulate the proposed concentrated particle-based receiver design three distinct but coupled physical phenomena must be modeled: fluid flows, Lagrangian particle dynamics, and the transport of thermal radiation. Therefore, three different physics solvers (fluid, particles, and radiation) must run concurrently with significant cross-communication in an integrated multi-physics simulation. However, each solver uses substantially different algorithms and data access patterns. Coordinating the overall data communication, computational load balancing, and scaling these different physics solvers together on modern massively parallel, heterogeneous high performance computing systems presents several major challenges. We have adopted the Legion programming system, via the Regent programming language, and its task parallel programming model to address these challenges. Our multi-physics solver Soleil-X is written entirely in the high level Regent programming language and is one of the largest and most complex applications written in Regent to date. At this workshop we will give an overview of the software architecture of Soleil-X as well as discuss how our multi-physics solver was designed to use the task parallel programming model provided by Legion. We will also discuss the development experience, scaling, performance, portability, and multi-physics simulation results.
Alexandros Nikolaos Ziogas, Tal Ben-Nun, Guillermo Indalecio Fernández, Timo Schneider, Mathieu Luisier, Torsten Hoefler
SC3
2019 Optimizing the data movement in quantum transport simulations via data-centric parallel programming
abstract
Designing efficient cooling systems for integrated circuits (ICs) relies on a deep understanding of the electro-thermal properties of transistors. To shed light on this issue in currently fabricated Fin-FETs, a quantum mechanical solver capable of revealing atomically-resolved electron and phonon transport phenomena from first-principles is required. In this paper, we consider a global, data-centric view of a state-of-the-art quantum transport simulator to optimize its execution on supercomputers. The approach yields coarse-and fine-grained data-movement characteristics, which are used for performance and communication modeling, communication-avoidance, and data-layout transformations. The transformations are tuned for the Piz Daint and Summit supercomputers, where each platform requires different caching and fusion strategies to perform optimally. The presented results make ab initio device simulation enter a new era, where nanostructures composed of over 10,000 atoms can be investigated at an unprecedented level of accuracy, paving the way for better heat management in next-generation ICs.
Alexandros Nikolaos Ziogas, Tal Ben-Nun, Guillermo Indalecio Fernández, Timo Schneider, Mathieu Luisier, Torsten Hoefler
SC3
2019 AXC: A new format to perform the SpMV oriented to Intel Xeon Phi architecture in OpenCL
abstract
Summary Emerging new architectures used in High Performance Computing require new research to adapt and optimise algorithms to them. As part of this effort, we propose the new AXC format to improve the performance of the SpMV product for the Intel Xeon Phi coprocessor. The performance of the OpenCL kernel, based on our new format, is compared with three very different and high efficient sparse matrix formats, ie, CSR, ELLR‐T, and K1. We perform tests with most of the matrices from the Williams collection used to test SpMV kernels for GPUs architectures in several related works. The numerical results show that the AXC format is more robust to spatial indirections proper of sparse matrices and prefers matrices with low variability amongst their rows' population, very much like matrices originated by FEM codes. The Conjugate Gradient (CG) is implemented in OpenCL using all the formats in this work to expose strengths and weaknesses of the formats in a real application. The CG implementation shows that the AXC has the fastest conversion time and its coherent with the numerical results generated by the SpMV tests, and that the format has a slower memory operations time due to an extra step required by the format and its larger memory footprint.
Edoardo Coronado-Barrientos, Guillermo Indalecio Fernández, Antonio J. García-Loureiro
Concurr. Comput. Pract. Exp.2
2018 Improving performance of iterative solvers with the AXC format using the Intel Xeon Phi
Edoardo Coronado-Barrientos, Guillermo Indalecio Fernández, Antonio J. García-Loureiro
J. Supercomput.2