EDBT 2026 Demo / reviewers in the wild / expert
Akshay Katta
dblp:40/6241
· DBLP profile ↗
2ranked-venue papers
0as first author
0since 2021 · last 2008
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Storage systems · 91% Distributed systems · 9% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems › data placement
data placement optimization |
0.1 | 1 | 2008 | Storage optimization for large-scale distributed stream-processing systems · ACM Trans. Storage 2008 |
Storage systems
distributed storage |
0.1 | 1 | 2008 | Storage optimization for large-scale distributed stream-processing systems · ACM Trans. Storage 2008 |
Storage systems › storage management
storage reclamation |
0.1 | 1 | 2008 | Storage optimization for large-scale distributed stream-processing systems · ACM Trans. Storage 2008 |
Distributed systems › stream processing
large-scale stream processing |
0.0 | 1 | 2008 | Storage optimization for large-scale distributed stream-processing systems · ACM Trans. Storage 2008 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.1optimization · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2008 | Storage optimization for large-scale distributed stream-processing systemsabstractWe consider storage in an extremely large-scale distributed computer system designed for stream processing applications. In such systems, both incoming data and intermediate results may need to be stored to enable analyses at unknown future times. The quantity of data of potential use would dominate even the largest storage system. Thus, a mechanism is needed to keep the data most likely to be used. One recently introduced approach is to employ retention value functions, which effectively assign each data object a value that changes over time in a prespecified way [Douglis et al.2004]. Storage space for data entering the system is reclaimed automatically by deleting data of the lowest current value. In such large systems, there will naturally be multiple file systems available, each with different properties. Choosing the right file system for a given incoming stream of data presents a challenge. In this article we provide a novel and effective scheme for optimizing the placement of data within a distributed storage subsystem employing retention value functions. The goal is to keep the data of highest overall value, while simultaneously balancing the read load to the file system. The key aspects of such a scheme are quite different from those that arise in traditional file assignment problems. We further motivate this optimization problem and describe a solution, comparing its performance to other reasonable schemes via simulation experiments. Kirsten Hildrum, Fred Douglis, Joel L. Wolf, Philip S. Yu, Lisa Fleischer, Akshay Katta |
ACM Trans. Storage | 6 |
| 2007 | Storage Optimization for Large-Scale Distributed Stream Processing SystemsabstractWe consider storage in an extremely large-scale distributed computer system designed for stream processing applications. In such systems, incoming data and intermediate results may need to be stored to enable future analyses. The quantity of such data would dominate even the largest storage system. Thus, a mechanism is needed to keep the most useful data. One recently introduced approach is to employ retention value functions, which effectively assign each data object a value that changes over time. Storage space is then reclaimed automatically by deleting data of lowest current value. In such large systems, there can naturally be multiple file systems available, each with different properties. Choosing the right file system for a given incoming data stream presents a challenge. In this paper we provide a novel and effective scheme for optimizing the placement of data within a distributed storage subsystem employing retention value functions. The goal is to keep the data of highest overall value, while simultaneously balancing the read load to the file system. Kirsten Hildrum, Fred Douglis, Joel L. Wolf, Philip S. Yu, Lisa Fleischer, Akshay Katta |
IPDPS | 6 |