Edward F. Valeev

dblp:46/2149 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
4since 2021 · last 2024
0000-0001-9923-6256ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 since 2021
YearPublicationVenuePosition
2024 CoNST: Code Generator for Sparse Tensor Networks
abstract
Sparse tensor networks represent contractions over multiple sparse tensors. Tensor contractions are higher-order analogs of matrix multiplication. Tensor networks arise commonly in many domains of scientific computing and data science. Such networks are typically computed using a tree of binary contractions. Several critical inter-dependent aspects must be considered in the generation of efficient code for a contraction tree, including sparse tensor layout mode order, loop fusion to reduce intermediate tensors, and the mutual dependence of loop order, mode order, and contraction order. We propose CoNST, a novel approach that considers these factors in an integrated manner using a single formulation. Our approach creates a constraint system that encodes these decisions and their interdependence, while aiming to produce reduced-order intermediate tensors via fusion. The constraint system is solved by the Z3 SMT solver and the result is used to create the desired fused loop structure and tensor mode layouts for the entire contraction tree. This structure is lowered to the IR of the TACO compiler, which is then used to generate executable code. Our experimental evaluation demonstrates significant performance improvements over current state-of-the-art sparse tensor compiler/library alternatives.
Saurabh Raje, Yufan Xu 0001, Atanas Rountev, Edward F. Valeev, P. Sadayappan
ACM Trans. Archit. Code Optim.4
2022 Pushing the Boundaries of Small Tasks: Scalable Low-Overhead Data-Flow Programming in TTG
abstract
Shared memory parallel programming models strive to provide low-overhead execution environments. Task-based programming models, in particular, are well-suited to cope with the ubiquitous multi- and many-core systems since they allow applications to express all available concurrency to a scheduler, which is tasked with exploiting the available hardware resources. It is general consensus that atomic operations should be preferred over locks and mutexes to avoid inter-thread serialization and the resulting loss in efficiency. However, even atomic operations may serialize threads if not used judiciously. In this work, we will discuss several optimizations applied to TTG and the underlying PaRSEC runtime system aiming at removing contentious atomic operations to reduce the overhead of task management to a few hundred clock cycles. The result is an optimized data-flow programming system that seamlessly scales from a single node to distributed execution and which is able to compete with OpenMP in shared memory.
Joseph Schuchart, Poornima Nookala, Thomas Hérault, Edward F. Valeev, George Bosilca
CLUSTER4
2022 Generalized Flow-Graph Programming Using Template Task-Graphs: Initial Implementation and Assessment
abstract
We present and evaluate TTG, a novel programming model and its C++ implementation that by marrying the ideas of control and data flowgraph programming supports compact specification and efficient distributed execution of dynamic and irregular applications. Programming interfaces that support task-based execution often only support shared memory parallel environments; a few support distributed memory environments, either by discovering the entire DAG of tasks on all processes, or by introducing explicit communications. The first approach limits scalability, while the second increases the complexity of programming. We demonstrate how TTG can address these issues without sacrificing scalability or programmability by providing higher-level abstractions than conventionally provided by task-centric programming systems, without impeding the ability of these runtimes to manage task creation and execution as well as data and resource management efficiently. TTG supports distributed memory execution over 2 different task runtimes, PaRSEC and MADNESS. Performance of four paradigmatic applications (in graph analytics, dense and block-sparse linear algebra, and numerical integrodifferential calculus) with various degrees of irregularity implemented in TTG is illustrated on large distributed-memory platforms and compared to the state-of-the-art implementations.
Joseph Schuchart, Poornima Nookala, Mohammad Mahdi Javanmard, Thomas Hérault, Edward F. Valeev, George Bosilca, Robert J. Harrison
IPDPS5
2021 Distributed-memory multi-GPU block-sparse tensor contraction for electronic structure
abstract
Many domains of scientific simulation (chemistry, condensed matter physics, data science) increasingly eschew dense tensors for block-sparse tensors, sometimes with additional structure (recursive hierarchy, rank sparsity, etc.). Distributed-memory parallel computation with block-sparse tensorial data is paramount to minimize the time-to-solution (e.g., to study dynamical problems or for real-time analysis) and to accommodate problems of realistic size that are too large to fit into the host/device memory of a single node equipped with accelerators. Unfortunately, computation with such irregular data structures is a poor match to the dominant imperative, bulk-synchronous parallel programming model. In this paper, we focus on the critical element of block-sparse tensor algebra, namely binary tensor contraction, and report on an efficient and scalable implementation using the task-focused PaRSEC runtime. High performance of the block-sparse tensor contraction on the Summit supercomputer is demonstrated for synthetic data as well as for real data involved in electronic structure simulations of unprecedented size.
Thomas Hérault, Yves Robert, George Bosilca, Robert J. Harrison, Cannada A. Lewis, Edward F. Valeev, Jack J. Dongarra
IPDPS6
2006 Poster reception - Component architectures for quantum chemistry: forging new capabilities and insights
abstract
We review the use of the Common Component Architecture approach within the quantum chemistry domain to tackle the software engineering challenges which arise as advanced algorithms are adopted and growing numbers of software packages are integrated to study complex, coupled physical phenomena. The development of common interfaces has allowed the adoption of advanced optimization solvers and high-level interchangeability of quantum chemistry packages. Components have been created which manage multiple levels of parallelism, providing much more efficient usage of parallel machines. Early efforts towards low-level integration of chemistry packages are examined. The ability to share intermediate data expands the capabilities available to any one software package, thereby enabling the rapid development of advanced methods. New methods for the study of reactions involving heavy elements, which depend on our component environment, are highlighted.
Joseph P. Kenny, Curtis L. Janssen, Ida M. B. Nielsen, Manojkumar Krishnan, Vidhya Gurumoorthi, Edward F. Valeev, Theresa L. Windus
SC6