VLDB 2026 Research / reviewers in the wild / expert
Eduardo F. D'Azevedo
dblp:81/7021
· DBLP profile ↗
9ranked-venue papers
1as first author
0since 2021 · last 2019
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 first-authorSoftware engineering, systems software and programming languages · 1Theory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
High-performance computing · 67% Emerging computing paradigms · 33% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational science and engineering · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing
performance optimization at scale |
0.1 | 1 | 2008 | New algorithm to enable 400+ TFlop/s sustained performance in simulations of disorder effects in high-Tc superconductors · SC 2008 |
Emerging computing paradigms › quantum computing › quantum simulation
quantum monte carlo simulation |
0.1 | 1 | 2008 | New algorithm to enable 400+ TFlop/s sustained performance in simulations of disorder effects in high-Tc superconductors · SC 2008 |
High-performance computing
scientific computing systems |
0.1 | 1 | 2008 | New algorithm to enable 400+ TFlop/s sustained performance in simulations of disorder effects in high-Tc superconductors · SC 2008 |
Computational science and engineering
computational physics |
0.0 | 1 | 2008 | New algorithm to enable 400+ TFlop/s sustained performance in simulations of disorder effects in high-Tc superconductors · SC 2008 |
Methods — techniques the papers use, named apart from their topics
mixed single-/double precision · 0.2delayed monte carlo updates · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Accelerating DCA++ (Dynamical Cluster Approximation) Scientific Application on the Summit SupercomputerabstractOptimizing scientific applications on today's accelerator-based high performance computing systems can be challenging, especially when multiple GPUs and CPUs with heterogeneous memories and persistent non-volatile memories are present. An example is Summit, an accelerator-based system at the Oak Ridge Leadership Computing Facility (OLCF) that is rated as the world's fastest supercomputer to-date. New strategies are thus needed to expose the parallelism in legacy applications, while being amenable to efficient mapping to the underlying architecture. In this paper we discuss our experiences and strategies to port a scientific application, DCA++, to Summit. DCA++ is a high-performance research application that solves quantum many-body problems with a cutting edge quantum cluster algorithm, the dynamical cluster approximation. Our strategies aim to synergize the strengths of the different programming models in the code. These include: a) streamlining the interactions between the CPU threads and the GPUs, b) implementing computing kernels on the GPUs and decreasing CPU-GPU memory transfers, c) allowing asynchronous GPU communications, and d) increasing compute intensity by combining linear algebraic operations. Full-scale production runs using all 4600 Summit nodes attained a peak performance of 73.5 PFLOPS with a mixed precision implementation. We observed a perfect strong and weak scaling for the quantum Monte Carlo solver in DCA++, while encountering about 2x input/output (I/O) and MPI communication overhead on the time-to-solution for the full machine run. Our hardware agnostic optimizations are designed to alleviate the communication and I/O challenges observed, while improving the compute intensity and obtaining optimal performance on a complex, hybrid architecture like Summit. Giovanni Balduzzi, Arghya Chatterjee 0001, Ying Wai Li, Peter W. Doak, Urs R. Hähner, Eduardo F. D'Azevedo, Thomas A. Maier, Thomas C. Schulthess |
PACT | 6 |
| 2018 | MiniApp for Density Matrix Renormalization Group Hamiltonian Application KernelabstractWe present two miniapps that implement the core computational kernel of the DMRG++ application, a generic C++ code that implements the Density Matrix Renormalization Group (DMRG) algorithm. The DMRG++ core Kronecker multiplication kernel is formulated using a batched BLAS approach, with implementation that targets both multi-core CPUs using OpenMP and GPGPU using the MAGMA library. The kernel evaluates the matrix-vector multiplication of the target Hamiltonian matrix used in Lanczos algorithm for computing the lowest eigenvalue and eigenvector. The Hamiltonian matrix is expressed compactly as sums of Kronecker products of small dense matrices. We demonstrate improved performance of the miniapp on synthetic problem, and show the performance of the DMRG++ application using a plugin based on the miniapp. We also present an OpenMP miniapp that explores the use of nested parallel constructs to implement the Kronecker multiplication kernel, exploring the use of nested OpenMP worksharing and tasking abstractions to implement the multi-level parallel multiplication algorithm. The miniapp has been used as a co-design vehicle for evaluating features in the OpenMP-4.5 and upcoming OpenMP-5.0 standards. Wael R. Elwasif, Eduardo F. D'Azevedo, Arghya Chatterjee 0001, Gonzalo Alvarez 0001, Oscar R. Hernandez, Vivek Sarkar |
CLUSTER | 2 |
| 2016 | Communication Characterization and Optimization of Applications Using Topology-Aware Task Mapping on Large SupercomputersabstractOn large supercomputers, the job scheduling systems may assign a non-contiguous node allocation for user applications depending on available resources. With parallel applications using MPI (Message Passing Interface), the default process ordering does not take into account the actual physical node layout available to the application. This contributes to non-locality in terms of physical network topology and impacts communication performance of the application. In order to mitigate such performance penalties, this work describes techniques to identify suitable task mapping that takes the layout of the allocated nodes as well as the application's communication behavior into account. During the first phase of this research, we instrumented and collected performance data to characterize communication behavior of critical US DOE (United States - Department of Energy) applications using an augmented version of the mpiP tool. Subsequently, we developed several reordering methods (spectral bisection, neighbor join tree etc.) to combine node layout and application communication data for optimized task placement. We developed a tool called mpiAproxy to facilitate detailed evaluation of the various reordering algorithms without requiring full application executions. This work presents a comprehensive performance evaluation (14,000 experiments) of the various task mapping techniques in lowering communication costs on Titan, the leadership class supercomputer at Oak Ridge National Laboratory. Sarat Sreepathi, Eduardo F. D'Azevedo, Bobby Philip, Patrick H. Worley |
ICPE | 2 |
| 2015 | Developing MiniApps on Modern Platforms Using Multiple Programming ModelsabstractWe have developed a set of reduced, proxy applications ("MiniApps") based on large-scale application codes supported at the Oak Ridge Leadership Computing Facility (OLCF). The MiniApps are designed to encapsulate the details of the most important (i.e. the most time-consuming and/or unique) facets of the applications that run in production mode on the OLCF. In each case, we have produced or plan to produce individual versions of the MiniApps using different specific programming models (e.g., OpenACC, CUDA, OpenMP). We describe some of our initial observations regarding these different implementations along with estimates of how closely the MiniApps track the actual performance characteristics (in particular, the overall scalability) of the large-scale applications from which they are derived. O. E. Bronson Messer, Eduardo F. D'Azevedo, Judith C. Hill, Wayne Joubert, S. Laosooksathit, Arnold N. Tharrington |
CLUSTER | 2 |
| 2011 | Efficient GPU Implementation for Particle in Cell AlgorithmabstractParticle in cell (PIC) algorithm is a widely used method in plasma physics to study the trajectories of charged particles under electromagnetic fields. The PIC algorithm is computationally intensive and its time requirements are proportional to the number of charged particles involved in the simulation. The focus of the paper is to parallelize the PIC algorithm on Graphics Processing Unit (GPU). We present several performance trade-offs related to small shared memory and atomic operations on the GPU to achieve high performance. Rejith George Joseph, Girish Ravunnikutty, Sanjay Ranka, Eduardo F. D'Azevedo, Scott Klasky |
IPDPS | 4 |
| 2010 | Complex version of high performance computing LINPACK benchmark (HPL)abstractAbstract This paper describes our effort to enhance the performance of the AORSA fusion energy simulation program through the use of high‐performance LINPACK (HPL) benchmark, commonly used in ranking the top 500 supercomputers. The algorithm used by HPL, enhanced by a set of tuning options, is more effective than that found in the ScaLAPACK library. Retrofitting these algorithms, such as look‐ahead processing of pivot elements, into ScaLAPACK is considered as a major undertaking. Moreover, HPL is configured as a benchmark, but only for real‐valued coefficients. We therefore developed software to convert HPL for use within an application program that generates complex coefficient linear systems. Although HPL is not normally perceived as a part of an application, our results show that the modified HPL software brings a significant increase in the performance of the solver when simulating the highest resolution experiments thus far configured, achieving 87.5 TFLOPS on over 20 000 processors on the Cray XT4. Copyright © 2009 John Wiley & Sons, Ltd. R. F. Barrett, T. H. F. Chan, Eduardo F. D'Azevedo, E. F. Jaeger, K. Wong, R. Y. Wong |
Concurr. Comput. Pract. Exp. | 3 |
| 2008 | New algorithm to enable 400+ TFlop/s sustained performance in simulations of disorder effects in high-Tc superconductorsabstractStaggering computational and algorithmic advances in recent years now make possible systematic Quantum Monte Carlo (QMC) simulations of high temperature (high-Tc) superconductivity in a microscopic model, the two dimensional (2D) Hubbard model, with parameters relevant to the cuprate materials. Here we report the algorithmic and computational advances that enable us to study the effect of disorder and nano-scale inhomogeneities on the pair-formation and the superconducting transition temperature necessary to understand real materials. The simulation code is written with a generic and extensible approach and is tuned to perform well at scale. Significant algorithmic improvements have been made to make effective use of current supercomputing architectures. By implementing delayed Monte Carlo updates and a mixed single-/double precision mode, we are able to dramatically increase the efficiency of the code. On the Cray XT4 systems of the Oak Ridge National Laboratory (ORNL), for example, we currently run production jobs on 31 thousand processors and thereby routinely achieve a sustained performance that exceeds 200 TFlop/s. On a system with 49 thousand processors we achieved a sustained performance of 409 TFlop/s. We present a study of how random disorder in the effective Coulomb interaction strength affects the superconducting transition temperature in the Hubbard model. Gonzalo Alvarez 0001, Michael S. Summers, Don E. Maxwell, Markus Eisenbach 0002, Jeremy S. Meredith, Jeffrey M. Larkin, John M. Levesque, Thomas A. Maier, Paul R. C. Kent, Eduardo F. D'Azevedo, Thomas C. Schulthess |
SC | 10 |
| 2008 | Algorithm 888: Spherical Harmonic Transform AlgorithmsabstractA collection of MATLAB classes for computing and using spherical harmonic transforms is presented. Methods of these classes compute differential operators on the sphere and are used to solve simple partial differential equations in a spherical geometry. The spectral synthesis and analysis algorithms using fast Fourier transforms and Legendre transforms with the associated Legendre functions are presented in detail. A set of methods associated with a spectral_field class provides spectral approximation to the differential operators ∇ ⋯, ∇ ×, ∇, and ∇ 2 in spherical geometry. Laplace inversion and Helmholtz equation solvers are also methods for this class. The use of the class and methods in MATLAB is demonstrated by the solution of the barotropic vorticity equation on the sphere. A survey of alternative algorithms is given and implementations for parallel high performance computers are discussed in the context of global climate and weather models. John B. Drake, Patrick H. Worley, Eduardo F. D'Azevedo |
ACM Trans. Math. Softw. | 3 |
| 2000 | The design and implementation of the parallel out-of-core ScaLAPACK LU, QR, and Cholesky factorization routinesabstractThis paper describes the design and implementation of three core factorization routines—LU, QR, and Cholesky—included in the out-of-core extension of ScaLAPACK. These routines allow the factorization and solution of a dense system that is too large to fit entirely in physical memory. The full matrix is stored on disk and the factorization routines transfer sub-matrice panels into memory. The ‘left-looking’ column-oriented variant of the factorization algorithm is implemented to reduce the disk I/O traffic. The routines are implemented using a portable I/O interface and utilize high-performance ScaLAPACK factorization routines as in-core computational kernels. We present the details of the implementation for the out-of-core ScaLAPACK factorization routines, as well as performance and scalability results on a Beowulf Linux cluster. Copyright © 2000 John Wiley & Sons, Ltd. Eduardo F. D'Azevedo, Jack J. Dongarra |
Concurr. Pract. Exp. | 1 |