VLDB 2026 Research / reviewers in the wild / expert
Tuomas Pelkonen
dblp:166/8373
· DBLP profile ↗
3ranked-venue papers
1as first author
1since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Cloud and datacenter computing · 69% Storage systems · 31% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Cloud and datacenter computing
cluster resource management and scheduling |
0.4 | 1 | 2020 | Twine: A Unified Cluster Management System for Shared Infrastructure · OSDI 2020 |
Cloud and datacenter computing
resource management |
0.4 | 1 | 2020 | Twine: A Unified Cluster Management System for Shared Infrastructure · OSDI 2020 |
Cloud and datacenter computing › cluster resource management and scheduling
geo-distributed scheduling |
0.2 | 1 | 2024 | MAST: Global Scheduling of ML Training across Geo-Distributed Datacenters at Hyperscale · OSDI 2024 |
Storage systems › data compression
time series compression |
0.2 | 1 | 2015 | Gorilla: A Fast, Scalable, In-Memory Time Series Database · Proc. VLDB Endow. 2015 |
Storage systems › data management
time series database |
0.2 | 1 | 2015 | Gorilla: A Fast, Scalable, In-Memory Time Series Database · Proc. VLDB Endow. 2015 |
Storage systems
storage reliability |
0.1 | 1 | 2015 | Gorilla: A Fast, Scalable, In-Memory Time Series Database · Proc. VLDB Endow. 2015 |
Methods — techniques the papers use, named apart from their topics
global scheduling · 1.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | MAST: Global Scheduling of ML Training across Geo-Distributed Datacenters at Hyperscale
Arnab Choudhury, Yang Wang 0009, Tuomas Pelkonen, Kutta Srinivasan, Abha Jain, Shenghao Lin, Delia David, Siavash Soleimanifard, Ritesh Tijoriwala, Denis Samoylov, Chunqiang Tang |
OSDI | 3 |
| 2020 | Twine: A Unified Cluster Management System for Shared Infrastructure
Chunqiang Tang, Kenny Yu, Kaushik Veeraraghavan, Jonathan Kaldor, Scott Michelson, Thawan Kooburat, Aravind Anbudurai, Kabir Gogia, Ben Christensen, Alex Gartrell, Maxim Khutornenko, Sachin Kulkarni, Marcin Pawlowski 0003, Tuomas Pelkonen, Andre Rodrigues, Rounak Tibrewal, Vaishnavi Venkatesan, Peter Zhang |
OSDI | 16 |
| 2015 | Gorilla: A Fast, Scalable, In-Memory Time Series DatabaseabstractLarge-scale internet services aim to remain highly available and responsive in the presence of unexpected failures. Providing this service often requires monitoring and analyzing tens of millions of measurements per second across a large number of systems, and one particularly effective solution is to store and query such measurements in a time series database (TSDB). A key challenge in the design of TSDBs is how to strike the right balance between efficiency, scalability, and reliability. In this paper we introduce Gorilla, Facebook's in-memory TSDB. Our insight is that users of monitoring systems do not place much emphasis on individual data points but rather on aggregate analysis, and recent data points are of much higher value than older points to quickly detect and diagnose the root cause of an ongoing problem. Gorilla optimizes for remaining highly available for writes and reads, even in the face of failures, at the expense of possibly dropping small amounts of data on the write path. To improve query efficiency, we aggressively leverage compression techniques such as delta-of-delta timestamps and XOR'd floating point values to reduce Gorilla's storage footprint by 10x. This allows us to store Gorilla's data in memory, reducing query latency by 73x and improving query throughput by 14x when compared to a traditional database (HBase)-backed time series data. This performance improvement has unlocked new monitoring and debugging tools, such as time series correlation search and more dense visualization tools. Gorilla also gracefully handles failures from a single-node to entire regions with little to no operational overhead. Tuomas Pelkonen, Scott Franklin 0002, Paul Cavallaro, Justin Meza, Justin Teller, Kaushik Veeraraghavan |
Proc. VLDB Endow. | 1 |