Kiril Dichev

dblp:05/8780 · DBLP profile ↗
← Back
9ranked-venue papers
8as first author
1since 2021 · last 2022
0000-0001-7817-5095ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 44% Energy-efficient computing · 44% Parallel and multicore computing · 13%
Computer networks
1 paper
Network measurement and analytics · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
fault tolerance
0.612022
Power Log'n'Roll: Power-Efficient Localized Rollback for MPI Applications Using Message Logging Protocols · IEEE Trans. Parallel Distributed Syst. 2022
Energy-efficient computing
power management
0.612022
Power Log'n'Roll: Power-Efficient Localized Rollback for MPI Applications Using Message Logging Protocols · IEEE Trans. Parallel Distributed Syst. 2022
Parallel and multicore computing › parallel programming models › message passing
MPI applications
0.212022
Power Log'n'Roll: Power-Efficient Localized Rollback for MPI Applications Using Message Logging Protocols · IEEE Trans. Parallel Distributed Syst. 2022
Network measurement and analytics › bandwidth estimation
available bandwidth estimation
0.112012
Efficient and reliable network tomography in heterogeneous networks using BitTorrent broadcasts and clustering algorithms · SC 2012
Network measurement and analytics
network tomography
0.112012
Efficient and reliable network tomography in heterogeneous networks using BitTorrent broadcasts and clustering algorithms · SC 2012
Network measurement and analytics
bandwidth estimation
0.012012
Efficient and reliable network tomography in heterogeneous networks using BitTorrent broadcasts and clustering algorithms · SC 2012

Methods — techniques the papers use, named apart from their topics

message logging · 0.6dynamic voltage and frequency scaling · 0.6clock modulation · 0.6clustering · 0.1bittorrent fragment counting · 0.1
YearPublicationVenuePosition
2022 Power Log'n'Roll: Power-Efficient Localized Rollback for MPI Applications Using Message Logging Protocols
abstract
In fault tolerance for parallel and distributed systems, message logging protocols have played a prominent role in the last three decades. Such protocols enable local rollback to provide recovery from fail-stop errors. Global rollback techniques can be straightforward to implement but at times lead to slower recovery than local rollback. Local rollback is more complicated but can offer faster recovery times. In this work, we study the power and energy efficiency implications of global and local rollback. We propose a power-efficient version of local rollback to reduce power consumption for non-critical,blockedprocesses, usingDynamic Voltage and Frequency Scaling(DVFS) andclock modulation(CM). Our results for 3 different MPI codes on 2 parallel systems show that power-efficient local rollback reduces CPU energy waste up to 50% during the recovery phase, compared to existing global and local rollback techniques, without introducing significant overheads. Furthermore, we show that savings manifest for all blocked processes, which grow linearly with the process count. We estimate that for settings with high recovery overheads the total energy waste of parallel codes is reduced with the proposed local rollback.
Kiril Dichev, Daniele De Sensi, Dimitrios S. Nikolopoulos, Kirk W. Cameron, Ivor T. A. Spence
IEEE Trans. Parallel Distributed Syst.1
2019 Implementing efficient message logging protocols as MPI application extensions
abstract
Message logging protocols are enablers of local rollback, a more efficient alternative to global rollback, for fault tolerant MPI applications. Until now, message logging MPI implementations have incurred the overheads of a redesign and redeployment of an MPI library, as well as continued performance penalties across various kernels. Successful research efforts for message logging implementations do exist, but not a single one of them can be easily deployed today by more than a few experts. In contrast, in this work we build efficient message logging capabilities on top of an MPI library with no message logging capabilities; we do so for two different send-deterministic HPC kernels, one with a global exchange pattern (CG), and one with a neighbour exchange pattern (LULESH). While our library of choice ULFM detects failure and recovers MPI communicators, we build on that to then restore the intra- and inter-process data consistency of both applications. This task follows a similar pattern across these kernels, and we present our methodology in a generic way. In the end, our extensions provide message logging capabilities for each kernel, without the need for an actual message logging runtime underneath. On the performance side, we eliminate event logging for these kernels, and design a flexible user-defined hybrid between global and local rollback. Our extensions span a few hundred lines of code for each kernel, are open-sourced, and enable local and global rollback after process failure.
Kiril Dichev, Dimitrios S. Nikolopoulos
EuroMPI1
2018 Energy-efficient localised rollback via data flow analysis and frequency scaling
abstract
Exascale systems will suffer failures hourly. HPC programmers rely mostly on application-level checkpoint and a global rollback to recover. In recent years, techniques reducing the number of rolling back processes have been implemented via message logging. However, the log-based approaches have weaknesses, such as being dependent on complex modifications within an MPI implementation, and the fact that a full restart may be required in the general case. To address the limitations of all log-based mechanisms, we return to checkpoint-only mechanisms, but advocate data flow rollback (DFR), a fundamentally different approach relying on analysis of the data flow of iterative codes, and the well-known concept of data flow graphs. We demonstrate the benefits of DFR for an MPI stencil code by localising rollback, and then reduce energy consumption by 10-12% on idling nodes via frequency scaling. We also provide large-scale estimates for the energy savings of DFR compared to global rollback, which for stencil codes increase as n2 for a process count n.
Kiril Dichev, Kirk W. Cameron, Dimitrios S. Nikolopoulos
EuroMPI1
2018 A taxonomy of task-based parallel programming technologies for high-performance computing
abstract
Task-based programming models for shared memory—such as Cilk Plus and OpenMP 3—are well established and documented. However, with the increase in parallel, many-core, and heterogeneous systems, a number of research-driven projects have developed more diversified task-based support, employing various programming and runtime features. Unfortunately, despite the fact that dozens of different task-based systems exist today and are actively used for parallel and high-performance computing (HPC), no comprehensive overview or classification of task-based technologies for HPC exists. In this paper, we provide an initial task-focused taxonomy for HPC technologies, which covers both programming interfaces and runtime mechanisms. We demonstrate the usefulness of our taxonomy by classifying state-of-the-art task-based environments in use today.
Peter Thoman, Kiril Dichev, Thomas Heller, Roman Iakymchuk, Xavier Aguilar, Khalid Hasanov, Philipp Gschwandtner, Pierre Lemarinier, Stefano Markidis, Herbert Jordan, Thomas Fahringer, Kostas Katrinis, Erwin Laure, Dimitrios S. Nikolopoulos
J. Supercomput.2
2016 TwinPCG: Dual Thread Redundancy with forward Recovery for Preconditioned Conjugate Gradient Methods
abstract
TwinPCG pays off where relatively high soft fault rates are anticipated. We see this possibility mainly in the area of nearthreshold voltage computing. In this domain, TwinPCG, and its dual redundancy, is a better choice than more resource-intensive strategies like TMR.
Kiril Dichev, Dimitrios S. Nikolopoulos
CLUSTER1
2016 TwinPCG: Dual Thread Redundancy with Forward Recovery for Preconditioned Conjugate Gradient Methods
abstract
Even though iterative solvers like the Preconditioned Conjugate Gradient method (PCG) have been studied for over fifty years, fault tolerance for such solvers has seen much attention in recent years. For iterative solvers, two major reliable strategies of recovery exist: checkpoint-restart for backward recovery, or some type of redundancy technique for forward recovery. Efficient low-overhead redundancy techniques like algorithm-based fault tolerance for sparse matrix-vector products (SpMxV) have recently been proposed. These techniques add resilience with a good, but limited scope; state-of-the-art techniques correct at most 1 fault within a SpMxV. In this work, we study a more powerful resilience concept, which is redundant multithreading. It offers more generic and stronger recovery guarantees, including any soft faults in PCG iterations (among others covering SpMxV), but also requires more resources. We carefully study this redundancy-efficiency conflict. We propose a fault-tolerant PCG method, called TwinPCG, which introduces very small wall-clock time overhead, and significant advantages in detection and correction strategies. Our method uses Dual Modular Redundancy instead of the more expensive Triple Modular Redundancy (TMR); still, it retains the TMR advantages of fault correction. We describe, implement, and benchmark our iterative solver, and compare it in terms of efficiency and fault tolerance capabilities to state-of-the-art techniques. We find that before multithreading in BLAS, TwinPCG introduces 5-6% runtime overhead compared to reference PCG implementations, and can exploit BLAS multithreading well. In the presence of faults, it reliably performs forward recovery for a range of problems, showing all the strengths of TMR techniques.
Kiril Dichev, Dimitrios S. Nikolopoulos
CLUSTER1
2012 Efficient and reliable network tomography in heterogeneous networks using BitTorrent broadcasts and clustering algorithms
abstract
In the area of network performance and discovery, network tomography focuses on reconstructing network properties using only end-to-end measurements at the application layer. One challenging problem in network tomography is reconstructing available bandwidth along all links during multiple source / multiple destination transmissions. The traditional measurement procedures used for bandwidth tomography are extremely time consuming. We propose a novel solution to this problem. Our method counts the fragments exchanged during a BitTorrent broadcast. While this measurement has a high level of randomness, it can be obtained very efficiently, and aggregated into a reliable metric. This data is then analyzed with state-of-the-art algorithms, which correctly reconstruct logical clusters of nodes interconnected by high bandwidth, as well as bottlenecks between these logical clusters. Our experiments demonstrate that the proposed two-phase approach efficiently solves the presented problem for a number of settings on a complex grid infrastructure.
Kiril Dichev, Fergal Reid, Alexey L. Lastovetsky
SC1
2011 Improvement of the Bandwidth of Cross-Site MPI Communication Using Optical Fiber
Kiril Dichev, Alexey L. Lastovetsky, Vladimir Rychkov
EuroMPI1
2010 Two Algorithms of Irregular Scatter/Gather Operations for Heterogeneous Platforms
Kiril Dichev, Vladimir Rychkov, Alexey L. Lastovetsky
EuroMPI1