EDBT 2026 Demo / reviewers in the wild / expert
Kassidy Barram
dblp:339/8340
· DBLP profile ↗
3ranked-venue papers in the field
1as first author
3since 2021 · last 2024
0009-0004-8999-5944ORCID · corroborated
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 3 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Scrybe: Enabling Programmatic Interfaces for Explorations Over Voluminous Spatiotemporal Data CollectionsabstractThis study focuses on enabling programmatic interfaces to perform exploratory analyses over voluminous data collections. The data we consider can be encoded in diverse formats and managed using diverse data storage frameworks. Our framework, code named Scrybe, manages the competing pulls of expressive computations and the need to manage resource utilization in shared clusters. The framework includes support for differentiated quality of service allowing preferentially higher resource utilization for certain users. We have validated our methodology with voluminous data collections housed in relational, NoSQL/document, and hybrid storage systems. Our performance benchmarks profile several aspects of our methodology, and demonstrate the effectiveness of our methodology. Kassidy Barram, Sangmi Lee Pallickara, Shrideep Pallickara |
BDCAT | 1 |
| 2023 | A Framework for Profiling Spatial Variability in the Performance of Classification ModelsabstractScientists use models to further their understanding of phenomena and inform decision-making. A confluence of factors has contributed to an exponential increase in spatial data volumes. In this study, we describe our methodology to identify spatial variation in the performance of classification models. Our methodology allows tracking a host of performance measures across different thresholds for the larger, encapsulating spatial area under consideration. Our methodology ensures frugal utilization of resources via a novel validation budgeting scheme that preferentially allocates observations for validations. We complement these efforts with a browser-based, GPU-accelerated visualization scheme that also incorporates support for streaming to assimilate validation results as they become available. Menuka Warushavithana, Kassidy Barram, Caleb Carlson, Saptashwa Mitra, Sudipto Ghosh 0001, F. Jay Breidt, Sangmi Lee Pallickara, Shrideep Pallickara |
BDCAT | 2 |
| 2022 | Resource Efficient Profiling of Spatial Variability in Performance of Regression ModelsabstractScientists design models to understand phenomena, make predictions, and/or inform decision-making. This study targets models that encapsulate spatially evolving phenomena. Given a model, our objective is to identify the accuracy of the model across all geospatial extents. A scientist may expect these validations to occur at varying spatial resolutions (e.g., states, counties, towns, and census tracts). Assessing a model with all available ground-truth data is infeasible due to the data volumes involved. We propose a framework to assess the performance of models at scale over diverse spatial data collections. Our methodology ensures orchestration of validation workloads while reducing memory strain, alleviating contention, enabling concurrency, and ensuring high throughput. We introduce the notion of a validation budget that represents an upper-bound on the total number of observations that are used to assess the performance of models across spatial extents. The validation budget attempts to capture the distribution characteristics of observations and is informed by multiple sampling strategies. Our design allows us to decouple the validation from the underlying model-fitting libraries to interoperate with models constructed using different libraries and analytical engines; our advanced research prototype currently supports Scikit-learn, PyTorch, and TensorFlow. Caleb Carlson, Menuka Warushavithana, Saptashwa Mitra, Kassidy Barram, Sudipto Ghosh 0001, F. Jay Breidt, Sangmi Lee Pallickara, Shrideep Pallickara |
IEEE Big Data | 4 |