VLDB 2026 Research / reviewers in the wild / expert
Eduardo Berrocal
dblp:141/5191
· DBLP profile ↗
6ranked-venue papers
4as first author
0since 2021 · last 2017
0000-0003-0570-7833ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 4 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Hardware reliability and fault tolerance · 50% Distributed systems · 39% Parallel and multicore computing · 4% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Distributed systems
fault tolerance |
0.5 | 2 | 2017 | Toward General Software Level Silent Data Corruption Detection for Parallel Applications · IEEE Trans. Parallel Distributed Syst. 2017 Lightweight Silent Data Corruption Detection Based on Runtime Data Analysis for HPC Applications · HPDC 2015 |
Hardware reliability and fault tolerance › error detection
silent data corruption detection |
0.5 | 2 | 2017 | Toward General Software Level Silent Data Corruption Detection for Parallel Applications · IEEE Trans. Parallel Distributed Syst. 2017 Lightweight Silent Data Corruption Detection Based on Runtime Data Analysis for HPC Applications · HPDC 2015 |
Distributed systems › replication
partial replication |
0.3 | 1 | 2017 | Toward General Software Level Silent Data Corruption Detection for Parallel Applications · IEEE Trans. Parallel Distributed Syst. 2017 |
Hardware reliability and fault tolerance › soft errors
silent data corruption |
0.3 | 1 | 2017 | Toward General Software Level Silent Data Corruption Detection for Parallel Applications · IEEE Trans. Parallel Distributed Syst. 2017 |
Hardware reliability and fault tolerance
soft errors |
0.2 | 1 | 2015 | Lightweight Silent Data Corruption Detection Based on Runtime Data Analysis for HPC Applications · HPDC 2015 |
Parallel and multicore computing › parallel programming models › message passing
MPI applications |
0.1 | 1 | 2017 | Toward General Software Level Silent Data Corruption Detection for Parallel Applications · IEEE Trans. Parallel Distributed Syst. 2017 |
High-performance computing
iterative applications |
0.1 | 1 | 2015 | Lightweight Silent Data Corruption Detection Based on Runtime Data Analysis for HPC Applications · HPDC 2015 |
Performance modeling and evaluation › statistical analysis › time series analysis
time series prediction |
0.1 | 1 | 2015 | Lightweight Silent Data Corruption Detection Based on Runtime Data Analysis for HPC Applications · HPDC 2015 |
Methods — techniques the papers use, named apart from their topics
replication · 0.3data-analytic prediction · 0.3time series prediction · 0.2runtime data analysis · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2017 | Topology mapping of irregular parallel applications on torus-connected supercomputers
Jingjin Wu, Xuanxing Xiong, Eduardo Berrocal, Zhiling Lan |
J. Supercomput. | 3 |
| 2017 | Toward General Software Level Silent Data Corruption Detection for Parallel ApplicationsabstractSilent data corruption (SDC) poses a great challenge for high-performance computing (HPC) applications as we move to extreme-scale systems. Mechanisms have been proposed that are able to detect SDC in HPC applications by using the peculiarities of the data (more specifically, its “smoothness” in time and space) to make predictions. However, these data-analytic solutions are still far from fully protecting applications to a level comparable with more expensive solutions such as full replication. In this work, we propose partial replication to overcome this limitation. More specifically, we have observed that not all processes of an MPI application experience the same level of data variability at exactly the same time. Thus, we can smartly choose and replicate only those processes for which the lightweight data-analytic detectors would perform poorly. In addition, we propose a new evaluation method based on the probability that a corruption will pass unnoticed by a particular detector (instead of just reporting overall single-bit precision and recall). In our experiments, we use four applications dealing with different explosions. Our results indicate that our new approach can protect the MPI applications analyzed with 7-70 percent less overhead (depending on the application) than that of full duplication with similar detection recall. Eduardo Berrocal, Leonardo Arturo Bautista-Gomez, Sheng Di, Zhiling Lan, Franck Cappello |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2016 | Exploring Partial Replication to Improve Lightweight Silent Data Corruption Detection for HPC Applications
Eduardo Berrocal, Leonardo Arturo Bautista-Gomez, Sheng Di, Zhiling Lan, Franck Cappello |
Euro-Par | 1 |
| 2015 | An Efficient Silent Data Corruption Detection Method with Error-Feedback Control and Even Sampling for HPC ApplicationsabstractThe silent data corruption (SDC) problem is attracting more and more attentions because it is expected to have a great impact on exascale HPC applications. SDC faults are hazardous in that they pass unnoticed by hardware and can lead to wrong computation results. In this work, we formulate SDC detection as a runtime one-step-ahead prediction method, leveraging multiple linear prediction methods in order to improve the detection results. The contributions are twofold: (1) we propose an error feedback control model that can reduce the prediction errors for different linear prediction methods, and (2) we propose a spatial-data-based even-sampling method to minimize the detection overheads (including memory and computation cost). We implement our algorithms in the fault tolerance interface, a fault tolerance library with multiple checkpoint levels, such that users can conveniently protect their HPC applications against both SDC errors and fail-stop errors. We evaluate our approach by using large-scale traces from well-known, large-scale HPC applications, as well as by running those HPC applications on a real cluster environment. Experiments show that our error feedback control model can improve detection sensitivity by 34-189% for bit-flip memory errors injected with the bit positions in the range [20,30], without any degradation on detection accuracy. Furthermore, memory size can be reduced by 33% with our spatial-data even-sampling method, with only a slight and graceful degradation in the detection sensitivity. Sheng Di, Eduardo Berrocal, Franck Cappello |
CCGRID | 2 |
| 2015 | Lightweight Silent Data Corruption Detection Based on Runtime Data Analysis for HPC ApplicationsabstractNext-generation supercomputers are expected to have more components and, at the same time, consume several times less energy per operation. Consequently, the number of soft errors is expected to increase dramatically in the coming years. In this respect, techniques that leverage certain properties of iterative HPC applications (such as the smoothness of the evolution of a particular dataset) can be used to detect silent errors at the application level. In this paper, we present a pointwise detection model with two phases: one involving the prediction of the next expected value in the time series for each data point, and another determining a range (i.e., normal value interval) surrounding the predicted next-step value. We show that dataset correlation can be used to detect corruptions indirectly and limit the size of the data set to monitor, taking advantage of the underlying physics of the simulation. Our results show that, using our techniques, we can detect a large number of corruptions (i.e., above 90% in some cases) with 84% memory overhead, and 13.75% extra computation time. Eduardo Berrocal, Leonardo Arturo Bautista-Gomez, Sheng Di, Zhiling Lan, Franck Cappello |
HPDC | 1 |
| 2014 | Exploring void search for fault detection on extreme scale systemsabstractMean Time Between Failures (MTBF), now calculated in days or hours, is expected to drop to minutes on exascale machines. The advancement of resilience technologies greatly depends on a deeper understanding of faults arising from hardware and software components. This understanding has the potential to help us build better fault tolerance technologies. For instance, it has been proved that combining checkpointing and failure prediction leads to longer checkpoint intervals, which in turn leads to fewer total checkpoints. In this paper we present a new approach for fault detection based on the Void Search (VS) algorithm. VS is used primarily in astrophysics for finding areas of space that have a very low density of galaxies. We evaluate our algorithm using real environmental logs from Mira Blue Gene/Q supercomputer at Argonne National Laboratory. Our experiments show that our approach can detect almost all faults (i.e., sensitivity close to 1) with a low false positive rate (i.e., specificity values above 0.7). We also compare our algorithm with a number of existing detection algorithms, and find that ours outperforms all of them. Eduardo Berrocal, Li Yu 0006, Sean Wallace, Michael E. Papka, Zhiling Lan |
CLUSTER | 1 |