Mauro Bianco

dblp:18/1689 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 2 first-authorDatabases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Memory systems · 40% Parallel and multicore computing · 21% Cloud and datacenter computing · 15%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 50% Compilers and program optimization · 50%

Topics — the 15 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.312018
RM-replay: a high-fidelity tuning, optimization and exploration tool for resource management · SC 2018
Memory systems
data locality
0.312017
Trends in Data Locality Abstractions for HPC Systems · IEEE Trans. Parallel Distributed Syst. 2017
Memory systems
data movement
0.312017
Trends in Data Locality Abstractions for HPC Systems · IEEE Trans. Parallel Distributed Syst. 2017
Memory systems › memory hierarchy
memory hierarchy management
0.312017
Trends in Data Locality Abstractions for HPC Systems · IEEE Trans. Parallel Distributed Syst. 2017
Program synthesis and code generation
domain-specific code generation
0.212015
STELLA: a domain-specific tool for structured grid methods in weather and climate models · SC 2015
Compilers and program optimization › loop optimization
stencil computation optimization
0.212015
STELLA: a domain-specific tool for structured grid methods in weather and climate models · SC 2015
High-performance computing
scientific computing systems
0.212015
STELLA: a domain-specific tool for structured grid methods in weather and climate models · SC 2015
Parallel and multicore computing
parallel programming models
0.222017
The STAPL parallel container framework · PPoPP 2011
Trends in Data Locality Abstractions for HPC Systems · IEEE Trans. Parallel Distributed Syst. 2017
Parallel and multicore computing
concurrent data structures
0.112011
The STAPL parallel container framework · PPoPP 2011
Distributed systems
distributed data structures
0.112011
The STAPL parallel container framework · PPoPP 2011
Performance modeling and evaluation › simulation
simulation-based evaluation
0.112018
RM-replay: a high-fidelity tuning, optimization and exploration tool for resource management · SC 2018
Parallel and multicore computing
parallel programming runtimes
0.112017
Trends in Data Locality Abstractions for HPC Systems · IEEE Trans. Parallel Distributed Syst. 2017
Environmental and earth informatics › atmospheric modeling
weather and climate modeling
0.112015
STELLA: a domain-specific tool for structured grid methods in weather and climate models · SC 2015
High-performance computing
many-core acceleration
0.112015
STELLA: a domain-specific tool for structured grid methods in weather and climate models · SC 2015
Parallel and multicore computing › parallel algorithms › parallel algorithm design
scalable parallel algorithms
0.012011
The STAPL parallel container framework · PPoPP 2011

Methods — techniques the papers use, named apart from their topics

domain-specific language · 0.7architecture-dependent code generation · 0.7replay-based tuning · 0.3optimization · 0.3survey · 0.3generic programming · 0.1composition · 0.1
YearPublicationVenuePosition
2018 Highly Scalable Stencil-Based Matrix-Free Stochastic Estimator for the Diagonal of the Inverse
abstract
Selected inversion problems must be addressed in several research fields like physics, genetics, weather forecasting, and finance, in order to extract selected entries from the inverse of large, sparse matrices. State-of-the-art algorithms are either based on the LU factorization or on an iterative process. Both approaches present computational bottlenecks related to prohibitive memory requirements or extremely high running time for large-scale matrices. In recent years, in order to overcome such limitations, an alternative approach for computing stochastic estimates of the inverse entries has been developed. In this work, we present a stochastic estimator for the diagonal of the inverse and test its performance on a dataset of symmetric, positive semidefinite matrices coming from the field of atomistic quantum transport simulations with nonequilibrium Green's functions (NEGF) formalism. In such a framework, it is required to solve the Schrödinger equation thousands of times, demanding the computation of the diagonal of the retarded Green's function, i.e., the inverse of a large, sparse matrix including open boundary conditions. Given the nature and the structure of the NEGF matrices, our stochastic estimation framework exploits the capabilities of a stencil-based, matrix-free code, avoiding the fill-in and lack of scalability that the LV-based methods present for three-dimensional nanoelectronic devices. We also illustrate the impact of the stochastic estimator by comparing its accuracy against existing methods and demonstrate its scalability performance on the “Piz Daint” cluster at the Swiss National Supercomputing Center, preparing for postpetascale three-dimensional nanoscale calculations.
Fabio Verbosio, Jurai Kardos, Mauro Bianco, Olaf Schenk
SBAC-PAD3
2018 RM-replay: a high-fidelity tuning, optimization and exploration tool for resource management
Maxime Martinasso, Miguel Gila, Mauro Bianco, Sadaf R. Alam, Colin McMurtrie, Thomas C. Schulthess
SC3
2017 Trends in Data Locality Abstractions for HPC Systems
abstract
The cost of data movement has always been an important concern in high performance computing (HPC) systems. It has now become the dominant factor in terms of both energy consumption and performance. Support for expression of data locality has been explored in the past, but those efforts have had only modest success in being adopted in HPC applications for various reasons. them However, with the increasing complexity of the memory hierarchy and higher parallelism in emerging HPC systems, locality management has acquired a new urgency. Developers can no longer limit themselves to low-level solutions and ignore the potential for productivity and performance portability obtained by using locality abstractions. Fortunately, the trend emerging in recent literature on the topic alleviates many of the concerns that got in the way of their adoption by application developers. Data locality abstractions are available in the forms of libraries, data structures, languages and runtime systems; a common theme is increasing productivity without sacrificing performance. This paper examines these trends and identifies commonalities that can combine various locality concepts to develop a comprehensive approach to expressing and managing data locality on future large-scale high-performance computing systems.
Didem Unat, Anshu Dubey, Torsten Hoefler, John Shalf, Mark James Abraham, Mauro Bianco, Bradford L. Chamberlain, Romain Cledat, H. Carter Edwards, Hal Finkel, Karl Fürlinger, Frank Hannig, Emmanuel Jeannot, Amir Kamil, Jeff Keasler, Paul H. J. Kelly, Vitus J. Leung, Hatem Ltaief, Naoya Maruyama, Chris J. Newburn, Miquel Pericàs
IEEE Trans. Parallel Distributed Syst.6
2015 STELLA: a domain-specific tool for structured grid methods in weather and climate models
abstract
Many high-performance computing applications solving partial differential equations (PDEs) can be attributed to the class of kernels using stencils on structured grids. Due to the disparity between floating point operation throughput and main memory bandwidth these codes typically achieve only a low fraction of peak performance. Unfortunately, stencil computation optimization techniques are often hardware dependent and lead to a significant increase in code complexity. We present a domain-specific tool, STELLA, which eases the burden of the application developer by separating the architecture dependent implementation strategy from the user-code and is targeted at multi- and manycore processors. On the example of a numerical weather prediction and regional climate model (COSMO) we demonstrate the usefulness of STELLA for a real-world production code. The dynamical core based on STELLA achieves a speedup factor of 1.8x (CPU) and 5.8x (GPU) with respect to the legacy code while reducing the complexity of the user code.
Tobias Gysi, Carlos Osuna, Oliver Fuhrer, Mauro Bianco, Thomas C. Schulthess
SC4
2014 A Generic Strategy for Multi-stage Stencils
Mauro Bianco, Ben Cumming
Euro-Par1
2011 The STAPL parallel container framework
abstract
The Standard Template Adaptive Parallel Library (STAPL) is a parallel programming infrastructure that extends C++ with support for parallelism. It includes a collection of distributed data structures called pContainers that are thread-safe, concurrent objects, i.e., shared objects that provide parallel methods that can be invoked concurrently. In this work, we present the STAPL Parallel Container Framework (PCF), that is designed to facilitate the development of generic parallel containers. We introduce a set of concepts and a methodology for assembling a pContainer from existing sequential or parallel containers, without requiring the programmer to deal with concurrency or data distribution issues. The PCF provides a large number of basic parallel data structures (e.g., pArray, pList, pVector, pMatrix, pGraph, pMap, pSet). The PCF provides a class hierarchy and a composition mechanism that allows users to extend and customize the current container base for improved application expressivity and performance. We evaluate STAPL pContainer performance on a CRAY XT4 massively parallel system and show that pContainer methods, generic pAlgorithms, and different applications provide good scalability on more than 16,000 processors.
Ilie Gabriel Tanase, Antal A. Buss, Adam Fidel, Harshvardhan, Ioannis Papadopoulos 0001, Olga Pearce, Timmie G. Smith, Nathan L. Thomas, Xiabing Xu, Nedal Mourad, Jeremy Vu, Mauro Bianco, Nancy M. Amato, Lawrence Rauchwerger
PPoPP12
2010 STAPL: standard template adaptive parallel library
abstract
The Standard Template Adaptive Parallel Library (stapl) is a high-productivity parallel programming framework that extends C++ and stl with unified support for shared and distributed memory parallelism. stapl provides distributed data structures (pContainers) and parallel algorithms (pAlgorithms) and a generic methodology for extending them to provide customized functionality. The stapl runtime system provides the abstraction for communication and program execution. In this paper, we describe the major components of stapl and present performance results for both algorithms and data structures showing scalability up to tens of thousands of processors.
Antal A. Buss, Harshvardhan, Ioannis Papadopoulos 0001, Olga Pearce, Timmie G. Smith, Ilie Gabriel Tanase, Nathan L. Thomas, Xiabing Xu, Mauro Bianco, Nancy M. Amato, Lawrence Rauchwerger
SYSTOR9
2009 Expectation of Strings with Mismatches under Markov Chain Distribution
Cinzia Pizzi, Mauro Bianco
SPIRE2
2007 Obtaining Performance Measures through Microbenchmarking in a Peer-to-Peer Overlay Computer
abstract
We address the problem of developing a suite of microbenchmarking experiments aimed at providing the basic functionalities of a measurement tool for a P2P-based globally distributed computing platform, usually referred to as overlay computer. We argue that such a measuring system should take into account the communication patterns generated by the applications in order to provide useful performance insights
Paolo Bertasi, Mauro Bianco, Andrea Pietracaprina, Geppino Pucci
CISIS2
2006 A Static Parallel Multifrontal Solver for Finite Element Meshes
Alberto Bertoldo, Mauro Bianco, Geppino Pucci
ISPA2
2003 A High-Performance UL Factorization for the Frontal Method
Mauro Bianco
ICCSA (1)1
2000 On the Predictive Quality of BSP-like Cost Functions for NOWs
Mauro Bianco, Geppino Pucci
Euro-Par1