EDBT 2026 Demo / reviewers in the wild / expert
Gregory R. Ganger
dblp:g/GregoryRGanger · also Greg Ganger
· DBLP profile ↗
26ranked-venue papers in the field
0as first author
1since 2021 · last 2024
0000-0002-3065-7316ORCID · verified
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 19Database Systems & Data Management · 6Data Mining & Knowledge Discovery · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Baleen: ML Admission & Prefetching for Flash Caches
Daniel Lin-Kit Wong, Carson Molder, Sathya Gunasekar, Jimmy Lu, Snehal Khandkar, Daniel S. Berger, Nathan Beckmann, Gregory R. Ganger |
FAST | 10 |
| 2019 | Cluster storage systems gotta have HeART: improving storage efficiency by exploiting disk-reliability heterogeneity
Saurabh Kadekodi, K. V. Rashmi, Gregory R. Ganger |
FAST | 3 |
| 2019 | Peering through the Dark: An Owl's View of Inter-job Dependencies and Jobs' Impact in Shared ClustersabstractShared multi-tenant infrastructures have enabled companies to consolidate workloads and data, increasing data-sharing and cross-organizational re-use of job outputs. This same resource- and work-sharing has also increased the risk of missed deadlines and diverging priorities as recurring jobs and workflows developed by different teams evolve independently. To prevent incidental business disruptions, identifying and managing job dependencies with clarity becomes increasingly important. Owl is a cluster log analysis and visualization tool that (i) extracts and visualizes job dependencies derived from historical job telemetry and data provenance data sets, and (ii) introduces a novel job valuation algorithm estimating the impact of a job on dependent users and jobs. This demonstration showcases Owl's features that can help users identify critical job dependencies and quantify job importance based on jobs' impact. Carlo Curino, Subru Krishnan, Konstantinos Karanasos, Panagiotis Garefalakis, Gregory R. Ganger |
SIGMOD Conference | 6 |
| 2017 | Online Deduplication for DatabasesabstractdbDedup is a similarity-based deduplication scheme for on-line database management systems (DBMSs). Beyond block-level compression of individual database pages or operation log (oplog) messages, as used in today's DBMSs, dbDedup uses byte-level delta encoding of individual records within the database to achieve greater savings. dbDedup's single-pass encoding method can be integrated into the storage and logging components of a DBMS to provide two benefits: (1) reduced size of data stored on disk beyond what traditional compression schemes provide, and (2) reduced amount of data transmitted over the network for replication services. To evaluate our work, we implemented dbDedup in a distributed NoSQL DBMS and analyzed its properties using four real datasets. Our results show that dbDedup achieves up to 37x reduction in the storage size and replication traffic of the database on its own and up to 61x reduction when paired with the DBMS's block-level compression. dbDedup provides both benefits with negligible effect on DBMS throughput or client latency (average and tail). Lianghong Xu, Andrew Pavlo, Sudipta Sengupta, Gregory R. Ganger |
SIGMOD Conference | 4 |
| 2014 | Toward strong, usable access control for shared distributed data
Michelle L. Mazurek, William Melicher, Manya Sleeper, Lujo Bauer, Gregory R. Ganger, Nitin Gupta 0001, Michael K. Reiter |
FAST | 6 |
| 2014 | SpringFS: bridging agility and performance in elastic distributed storage
Lianghong Xu, James Cipar, Elie Krevat, Alexey Tumanov, Nitin Gupta 0001, Michael A. Kozuch, Gregory R. Ganger |
FAST | 7 |
| 2012 | RainMon: an integrated approach to mining bursty timeseries monitoring dataabstractMetrics like disk activity and network traffic are widespread sources of diagnosis and monitoring information in datacenters and networks. However, as the scale of these systems increases, examining the raw data yields diminishing insight. We present RainMon, a novel end-to-end approach for mining timeseries monitoring data designed to handle its size and unique characteristics. Our system is able to (a) mine large, bursty, real-world monitoring data, (b) find significant trends and anomalies in the data, (c) compress the raw data effectively, and (d) estimate trends to make forecasts. Furthermore, RainMon integrates the full analysis process from data storage to the user interface to provide accessible long-term diagnosis. We apply RainMon to three real-world datasets from production systems and show its utility in discovering anomalous machines and time periods. Ilari Shafer, Kai Ren 0001, Vishnu Naresh Boddeti, Yoshihisa Abe, Gregory R. Ganger, Christos Faloutsos |
KDD | 5 |
| 2009 | Perspective: Semantic Data Management for the Home
Brandon Salmon, Steven W. Schlosser, Lorrie Faith Cranor, Gregory R. Ganger |
FAST | 4 |
| 2008 | Measurement and Analysis of TCP Throughput Collapse in Cluster-based Storage Systems
Amar Phanishayee, Elie Krevat, Vijay Vasudevan, David G. Andersen, Gregory R. Ganger, Garth A. Gibson, Srinivasan Seshan |
FAST | 5 |
| 2008 | Using Utility to Provision Storage Systems
John D. Strunk, Eno Thereska, Christos Faloutsos, Gregory R. Ganger |
FAST | 4 |
| 2007 | //TRACE: Parallel Trace Replay with Approximate Causal Events
Michael P. Mesnier, Matthew Wachs, Raja R. Sambasivan, Julio López 0002, James Hendricks, Gregory R. Ganger, David R. O'Hallaron |
FAST | 6 |
| 2007 | Argon: Performance Insulation for Shared Storage Servers
Matthew Wachs, Michael Abd-El-Malek, Eno Thereska, Gregory R. Ganger |
FAST | 4 |
| 2007 | MultiMap: Preserving disk locality for multidimensional datasetsabstractMultiMap is an algorithm for mapping multidimensional datasets so as to preserve the data's spatial locality on disks. Without revealing disk-specific details to applications, MultiMap exploits modern disk characteristics to provide full streaming bandwidth for one (primary) dimension and maximally efficient non-sequential access (i.e., minimal seek and no rotational latency) for the other dimensions. This is in contrast to existing approaches, which either severely penalize non-primary dimensions or fail to provide full streaming bandwidth for any dimension. Experimental evaluation of a prototype implementation demonstrates MultiMap's superior performance for range and beam queries. On average, MultiMap reduces total I/O time by over 50% when compared to traditional linearized layouts and by over 30% when compared to space-filling curve approaches such as Z-ordering and Hilbert curves. For scans of the primary dimension, MultiMap and traditional linearized layouts provide almost two orders of magnitude higher throughput than space-filling curve approaches. Minglong Shao, Steven W. Schlosser, Stratos Papadomanolakis, Jiri Schindler, Anastasia Ailamaki, Gregory R. Ganger |
ICDE | 6 |
| 2005 | Ursa Minor: Versatile Cluster-based Storage
Michael Abd-El-Malek, William V. Courtright II, Chuck Cranor, Gregory R. Ganger, James Hendricks, Andrew J. Klosterman, Michael P. Mesnier, Manish Prasad, Brandon Salmon, Raja R. Sambasivan, Shafeeq Sinnamohideen, John D. Strunk, Eno Thereska, Matthew Wachs, Jay J. Wylie |
FAST | 4 |
| 2005 | On Multidimensional Data and Modern Disks
Steven W. Schlosser, Jiri Schindler, Stratos Papadomanolakis, Minglong Shao, Anastasia Ailamaki, Christos Faloutsos, Gregory R. Ganger |
FAST | 7 |
| 2004 | Diamond: A Storage Architecture for Early Discard in Interactive Search
Larry Huston, Rahul Sukthankar, Rajiv Wickremesinghe, Mahadev Satyanarayanan, Gregory R. Ganger, Erik Riedel, Anastasia Ailamaki |
FAST | 5 |
| 2004 | Atropos: A Disk Array Volume Manager for Orchestrated Use of Disks
Jiri Schindler, Steven W. Schlosser, Minglong Shao, Anastasia Ailamaki, Gregory R. Ganger |
FAST | 5 |
| 2004 | MEMS-based Storage Devices and Standard Disk Interfaces: A Square Peg in a Round Hole?
Steven W. Schlosser, Gregory R. Ganger |
FAST | 2 |
| 2004 | A Framework for Building Unobtrusive Disk Maintenance Applications (Awarded Best Student Paper!)
Eno Thereska, Jiri Schindler, John S. Bucy, Brandon Salmon, Christopher R. Lumb, Gregory R. Ganger |
FAST | 6 |
| 2004 | Clotho: Decoupling memory page layout from storage organization
Minglong Shao, Jiri Schindler, Steven W. Schlosser, Anastasia Ailamaki, Gregory R. Ganger |
VLDB | 5 |
| 2003 | Metadata Efficiency in Versioning File Systems
Craig A. N. Soules, Garth R. Goodson, John D. Strunk, Gregory R. Ganger |
FAST | 4 |
| 2003 | Lachesis: Robust Database Storage Management Based on Device-specific Performance Characteristics
Jiri Schindler, Anastasia Ailamaki, Gregory R. Ganger |
VLDB | 3 |
| 2002 | Timing-Accurate Storage Emulation
John Linwood Griffin, Jiri Schindler, Steven W. Schlosser, John S. Bucy, Gregory R. Ganger |
FAST | 5 |
| 2002 | Freeblock Scheduling Outside of Disk Firmware
Christopher R. Lumb, Jiri Schindler, Gregory R. Ganger |
FAST | 3 |
| 2002 | Track-Aligned Extents: Matching Access Patterns to Disk Drive Characteristics
Jiri Schindler, John Linwood Griffin, Christopher R. Lumb, Gregory R. Ganger |
FAST | 4 |
| 2000 | Data Mining on an OLTP System (Nearly) for FreeabstractThis paper proposes a scheme for scheduling disk requests that takes advantage of the ability of high-level functions to operate directly at individual disk drives. We show that such a scheme makes it possible to support a Data Mining workload on an OLTP system almost for free: there is only a small impact on the throughput and response time of the existing workload. Specifically, we show that an OLTP system has the disk resources to consistently provide one third of its sequential bandwidth to a background Data Mining task with close to zero impact on OLTP throughput and response time at high transaction loads. At low transaction loads, we show much lower impact than observed in previous work. This means that a production OLTP system can be used for Data Mining tasks without the expense of a second dedicated system. Our scheme takes advantage of close interaction with the on-disk scheduler by reading blocks for the Data Mining workload as the disk head “passes over” them while satisfying demand blocks from the OLTP request stream. We show that this scheme provides a consistent level of throughput for the background workload even at very high foreground loads. Such a scheme is of most benefit in combination with an Active Disk environment that allows the background Data Mining application to also take advantage of the processing power and memory available directly on the disk drives. Erik Riedel, Christos Faloutsos, Gregory R. Ganger, David Nagle |
SIGMOD Conference | 3 |