VLDB 2026 Research / reviewers in the wild / expert
Carlos H. A. Costa
dblp:117/6175
· DBLP profile ↗
8ranked-venue papers
2as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-authorComputer networks · 2Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
High-performance computing · 33% Storage systems · 17% Distributed systems · 14% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 15 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › systems biology
multiscale modeling |
0.4 | 1 | 2019 | A massively parallel infrastructure for adaptive multiscale simulations: modeling RAS initiation pathway for cancer · SC 2019 |
High-performance computing
application porting |
0.4 | 1 | 2019 | Preparation and optimization of a diverse workload for a large-scale heterogeneous system · SC 2019 |
Parallel and multicore computing
programming models |
0.4 | 1 | 2019 | Preparation and optimization of a diverse workload for a large-scale heterogeneous system · SC 2019 |
High-performance computing
scientific computing systems |
0.4 | 1 | 2019 | A massively parallel infrastructure for adaptive multiscale simulations: modeling RAS initiation pathway for cancer · SC 2019 |
Storage systems
adaptive prefetching |
0.3 | 1 | 2017 | Leveraging Adaptive I/O to Optimize Collective Data Shuffling Patterns for Big Data Analytics · IEEE Trans. Parallel Distributed Syst. 2017 |
Cloud and datacenter computing
big data analytics |
0.3 | 1 | 2017 | Leveraging Adaptive I/O to Optimize Collective Data Shuffling Patterns for Big Data Analytics · IEEE Trans. Parallel Distributed Syst. 2017 |
Distributed systems › distributed data processing
data shuffling |
0.3 | 1 | 2017 | Leveraging Adaptive I/O to Optimize Collective Data Shuffling Patterns for Big Data Analytics · IEEE Trans. Parallel Distributed Syst. 2017 |
Storage systems
i/o optimization |
0.3 | 1 | 2017 | Leveraging Adaptive I/O to Optimize Collective Data Shuffling Patterns for Big Data Analytics · IEEE Trans. Parallel Distributed Syst. 2017 |
Hardware reliability and fault tolerance › memory reliability
DRAM errors |
0.2 | 1 | 2014 | A System Software Approach to Proactive Memory-Error Avoidance · SC 2014 |
Distributed systems
fault tolerance |
0.2 | 1 | 2014 | A System Software Approach to Proactive Memory-Error Avoidance · SC 2014 |
Memory systems › virtual memory management
page migration |
0.2 | 1 | 2014 | A System Software Approach to Proactive Memory-Error Avoidance · SC 2014 |
Bioinformatics and computational biology › computational oncology
cancer modeling |
0.1 | 1 | 2019 | A massively parallel infrastructure for adaptive multiscale simulations: modeling RAS initiation pathway for cancer · SC 2019 |
Performance modeling and evaluation
benchmarking |
0.1 | 1 | 2019 | Preparation and optimization of a diverse workload for a large-scale heterogeneous system · SC 2019 |
Cloud and datacenter computing › big data analytics
in-memory data analytics |
0.1 | 1 | 2017 | Leveraging Adaptive I/O to Optimize Collective Data Shuffling Patterns for Big Data Analytics · IEEE Trans. Parallel Distributed Syst. 2017 |
Operating systems › resource management
memory management |
0.1 | 1 | 2014 | A System Software Approach to Proactive Memory-Error Avoidance · SC 2014 |
Methods — techniques the papers use, named apart from their topics
log analysis · 0.4error prediction · 0.4prefetching · 0.3adaptive i/o · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | You Only Run Once: Spark Auto-Tuning From a Single RunabstractTuning configurations of Spark jobs is not a trivial task. State-of-the-art auto-tuning systems are based on iteratively running workloads with different configurations. During the optimization process, the relevant features are explored to find good solutions. Many optimizers enhance the time-to-solution using black-box optimization algorithms that do not take into account any information from the Spark workloads. In this article, we present a new method for tuning configurations that uses information from one run of a Spark workload. To achieve good performance, we mine the SparkEventLog that is generated by the Spark engine. This log file contains a large amount of information from the executed application. We use this information to enhance a performance model with low-level features from the workload to be optimized. These features include Spark Actions, Transformations, and Task metrics. This process allows us to obtain application-specific workload information. With this information our system can predict sensible Spark configurations for unseen jobs, given that it has been trained with reasonable coverage of Spark applications. Experiments show that the presented system correctly produces good configurations, while achieving up to 80% speedup with respect to the default Spark configuration, and up to 12x speedup of the time-to-solution with respect to a standard Bayesian Optimization procedure. David Buchaca Prats, Felipe Albuquerque Portella, Carlos H. A. Costa, Josep Lluís Berral |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2019 | Preparation and optimization of a diverse workload for a large-scale heterogeneous systemabstractProductivity from day one on supercomputers that leverage new technologies requires significant preparation. An institution that procures a novel system architecture often lacks sufficient institutional knowledge and skills to prepare for it. Thus, the "Center of Excellence" (CoE) concept has emerged to prepare for systems such as Summit and Sierra, currently the top two systems in the Top 500. This paper documents CoE experiences that prepared a workload of diverse applications and math libraries for a heterogeneous system. We describe our approach to this preparation, including our management and execution strategies, and detail our experiences with and reasons for using different programming approaches. Our early science and performance results show that the project enabled significant early seismic science with up to a l4X throughput increase over Cori. In addition to our successes, we discuss our challenges and failures so others may benefit from our experience. Ian Karlin, Yoonho Park, Bronis R. de Supinski, Bert Still, D. A. Beckingsale, Robert Blake, Tong Chen 0001, Guojing Cong, Carlos H. A. Costa, Johann Dahm, Giacomo Domeniconi, Thomas Epperly, Aaron Fisher, Sara Kokkila Schumacher, Steve H. Langer, Hai Le, Naoya Maruyama, Xinyu Que, David F. Richards, Björn Sjögreen, Jonathan Wong, Carol S. Woodward, Ulrike Meier Yang, Bob Anderson, David Appelhans, Levi Barnes, Peter D. Barnes Jr., Sorin Bastea, David Böhme, Jamie A. Bramwell, James M. Brase, José R. Brunheroto, Barry Chen, Charway R. Cooper, Tony Degroot, Robert D. Falgout, Todd Gamblin, David J. Gardner, James N. Glosli, John A. Gunnels, Max P. Katz, Tzanio V. Kolev, I-Feng W. Kuo, Matthew P. LeGendre, Pei-Hung Lin, Shelby Lockhart, Kathleen McCandless, Claudia Misale, Jaime H. Moreno, Rob Neely, Jarom Nelson, Rao Nimmakayala, Kathryn M. O'Brien, Kevin O'Brien, Ramesh Pankajakshan, Roger A. Pearce, Slaven Peles, Phil Regier, Steven C. Rennich, Martin Schulz 0001, Howard Scott, James C. Sexton, Kathleen Shoga, Shiv Sundram, Guillaume Thomas-Collignon, Brian Van Essen, Alexey Voronin, Bob Walkup, Chris Ward, Hui-Fang Wen, Daniel A. White, Christopher Young, Cyril Zeller, Edward Zywicz |
SC | 10 |
| 2019 | A massively parallel infrastructure for adaptive multiscale simulations: modeling RAS initiation pathway for cancerabstractComputational models can define the functional dynamics of complex systems in exceptional detail. However, many modeling studies face seemingly incommensurate requirements: to gain meaningful insights into some phenomena requires models with high resolution (microscopic) detail that must nevertheless evolve over large (macroscopic) length- and time-scales. Multiscale modeling has become increasingly important to bridge this gap. Executing complex multiscale models on current petascale computers with high levels of parallelism and heterogeneous architectures is challenging. Many distinct types of resources need to be simultaneously managed, such as GPUs and CPUs, memory size and latencies, communication bottlenecks, and filesystem bandwidth. In addition, robustness to failure of compute nodes, network, and filesystems is critical. Francesco Di Natale, Harsh Bhatia, Timothy S. Carpenter, Chris Neale, Sara Kokkila Schumacher, Tomas Oppelstrup, Liam Stanton, Shiv Sundram, Thomas Scogland, Gautham Dharuman, Michael P. Surh, Yue Yang 0034, Claudia Misale, Lars Schneidenbach, Carlos H. A. Costa, Changhoan Kim, Bruce D'Amora, Sandrasegaram Gnanakaran, Dwight V. Nissley, Frederick H. Streitz, Felice C. Lightstone, Peer-Timo Bremer, James N. Glosli, Helgi I. Ingólfsson |
SC | 16 |
| 2018 | Optimization of Genomics Analysis Pipeline for Scalable Performance in a Cloud Environment
Carlos H. A. Costa, Claudia Misale, Frank Liu 0001, Marcio Silva, Hubertus Franke, Paul Crumley, Bruce D'Amora |
BIBM | 1 |
| 2017 | Leveraging Adaptive I/O to Optimize Collective Data Shuffling Patterns for Big Data AnalyticsabstractBig data analytics is an indispensable tool in transforming science, engineering, medicine, health-care, finance and ultimately business itself. With the explosion of data sizes and need for shorter time-to-solution, in-memory platforms such as Apache Spark gain increasing popularity. In this context, data shuffling, a particularly difficult transformation pattern, introduces important challenges. Specifically, data shuffling is a key component of complex computations that has a major impact on the overall performance and scalability. Thus, speeding up data shuffling is a critical goal. To this end, state-of-the-art solutions often rely on overlapping the data transfers with the shuffling phase. However, they employ simple mechanisms to decide how much data and where to fetch it from, which leads to sub-optimal performance and excessive auxiliary memory utilization for the purpose of prefetching. The latter aspect is a growing concern, given evidence that memory per computation unit is continuously decreasing while interconnect bandwidth is increasing. This paper contributes a novel shuffle data transfer strategy that addresses the two aforementioned dimensions by dynamically adapting the prefetching to the computation. We implemented this novel strategy in Spark, a popular inmemory data analytics framework. To demonstrate the benefits of our proposal, we run extensive experiments on an HPC cluster with large core count per node. Compared with the default Spark shuffle strategy, our proposal shows: up to 40 percent better performance with 50 percent less memory utilization for buffering and excellent weak scalability. Bogdan Nicolae, Carlos H. A. Costa, Claudia Misale, Kostas Katrinis, Yoonho Park |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2016 | Towards Memory-Optimized Data Shuffling Patterns for Big Data AnalyticsabstractBig data analytics is an indispensable tool in transforming science, engineering, medicine, healthcare, finance and ultimately business itself. With the explosion of data sizes and need for shorter time-to-solution, in-memory platforms such as Apache Spark gain increasing popularity. However, this introduces important challenges, among which data shuffling is particularly difficult: on one hand it is a key part of the computation that has a major impact on the overall performance and scalability so its efficiency is paramount, while on the other hand it needs to operate with scarce memory in order to leave as much memory available for data caching. In this context, efficient scheduling of data transfers such that it addresses both dimensions of the problem simultaneously is non-trivial. State-of-the-art solutions often rely on simple approaches that yield sub optimal performance and resource usage. This paper contributes a novel shuffle data transfer strategy that dynamically adapts to the computation with minimal memory utilization, which we briefly underline as a series of design principles. Bogdan Nicolae, Carlos H. A. Costa, Claudia Misale, Kostas Katrinis, Yoonho Park |
CCGrid | 2 |
| 2014 | A System Software Approach to Proactive Memory-Error AvoidanceabstractToday's HPC systems use two mechanisms to address main-memory errors. Error-correcting codes make correctable errors transparent to software, while checkpoint/restart (CR) enables recovery from uncorrectable errors. Unfortunately, CR overhead will be enormous at exascale due to the high failure rate of memory. We propose a new OS-based approach that proactively avoids memory errors using prediction. This scheme exposes correctable error information to the OS, which migrates pages and off lines unhealthy memory to avoid application crashes. We analyze memory error patterns in extensive logs from a BG/P system and show how correctable error patterns can be used to identify memory likely to fail. We implement a proactive memory management system on BG/Q by extending the firmware and Linux. We evaluate our approach with a realistic workload and compare our overhead against CR. We show improved resilience with negligible performance overhead for applications. Carlos H. A. Costa, Yoonho Park, Bryan S. Rosenburg, Chen-Yong Cher, Kyung Dong Ryu |
SC | 1 |
| 2013 | Evaluation of a policy-based network management system for energy-efficiency
Guilherme Carvalho Januario, Carlos H. A. Costa, Marcelo C. Amarai, Ana C. Riekstin, Tereza Cristina M. B. Carvalho, Catalin Meirosu |
IM | 2 |