Alexander Balgavy

dblp:405/3138 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2026
0009-0009-0521-3404ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 44% Distributed systems · 44% Performance modeling and evaluation · 13%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
cloud service reliability
1.012026
Cloud Uptime Archive: Open-Access Availability Data of Web, Cloud, and Gaming Services · IEEE Trans. Parallel Distributed Syst. 2026
Distributed systems
fault tolerance
1.012026
Cloud Uptime Archive: Open-Access Availability Data of Web, Cloud, and Gaming Services · IEEE Trans. Parallel Distributed Syst. 2026
Performance modeling and evaluation › simulation
simulation-based evaluation
0.312026
Cloud Uptime Archive: Open-Access Availability Data of Web, Cloud, and Gaming Services · IEEE Trans. Parallel Distributed Syst. 2026

Methods — techniques the papers use, named apart from their topics

retry mechanisms · 1.0checkpointing simulation · 1.0MTBF/MTTR analysis · 1.0
YearPublicationVenuePosition
2026 Cloud Uptime Archive: Open-Access Availability Data of Web, Cloud, and Gaming Services
abstract
Cloud services are critical to society. However, their reliability is poorly understood. Towards solving the problem, we propose a standard repository for cloud uptime data. We populate this repository with the data we collect containing failure reports from users and operators of cloud services, web services, and online games. The multiple vantage points help reduce bias from individual users and operators. We compare our new data to existing failure data from the Failure Trace Archive and the Google cluster trace. We analyze the MTBF and MTTR, time patterns, failure severity, user-reported symptoms, and operator-reported symptoms of failures in the data we collect. We observe that high-level user facing services fail less often than low-level infrastructure services, likely due to them using fault-tolerance techniques. We use simulation-based experiments to demonstrate the impact of different failure traces on the performance of checkpointing and retry mechanisms. We release the data, and the analysis and simulation tools, as open-source artifacts available athttps://github.com/atlarge-research/cloud-uptime-archive.
Sacheendra Talluri, Dante Niewenhuis, Xiaoyu Chu, Jakob Kyselica, Mehmet Çetin, Alexander Balgavy, Alexandru Iosup
IEEE Trans. Parallel Distributed Syst.6