EDBT 2026 Demo / reviewers in the wild / expert
Sudipto Ghosh 0001
dblp:244/6021-1
· DBLP profile ↗
6ranked-venue papers in the field
0as first author
2since 2021 · last 2023
0000-0001-6000-9646ORCID · verified
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 4Database Systems & Data Management · 1Knowledge Engineering, Semantic Web & Information Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | A Framework for Profiling Spatial Variability in the Performance of Classification ModelsabstractScientists use models to further their understanding of phenomena and inform decision-making. A confluence of factors has contributed to an exponential increase in spatial data volumes. In this study, we describe our methodology to identify spatial variation in the performance of classification models. Our methodology allows tracking a host of performance measures across different thresholds for the larger, encapsulating spatial area under consideration. Our methodology ensures frugal utilization of resources via a novel validation budgeting scheme that preferentially allocates observations for validations. We complement these efforts with a browser-based, GPU-accelerated visualization scheme that also incorporates support for streaming to assimilate validation results as they become available. Menuka Warushavithana, Kassidy Barram, Caleb Carlson, Saptashwa Mitra, Sudipto Ghosh 0001, F. Jay Breidt, Sangmi Lee Pallickara, Shrideep Pallickara |
BDCAT | 5 |
| 2022 | Resource Efficient Profiling of Spatial Variability in Performance of Regression ModelsabstractScientists design models to understand phenomena, make predictions, and/or inform decision-making. This study targets models that encapsulate spatially evolving phenomena. Given a model, our objective is to identify the accuracy of the model across all geospatial extents. A scientist may expect these validations to occur at varying spatial resolutions (e.g., states, counties, towns, and census tracts). Assessing a model with all available ground-truth data is infeasible due to the data volumes involved. We propose a framework to assess the performance of models at scale over diverse spatial data collections. Our methodology ensures orchestration of validation workloads while reducing memory strain, alleviating contention, enabling concurrency, and ensuring high throughput. We introduce the notion of a validation budget that represents an upper-bound on the total number of observations that are used to assess the performance of models across spatial extents. The validation budget attempts to capture the distribution characteristics of observations and is informed by multiple sampling strategies. Our design allows us to decouple the validation from the underlying model-fitting libraries to interoperate with models constructed using different libraries and analytical engines; our advanced research prototype currently supports Scikit-learn, PyTorch, and TensorFlow. Caleb Carlson, Menuka Warushavithana, Saptashwa Mitra, Kassidy Barram, Sudipto Ghosh 0001, F. Jay Breidt, Sangmi Lee Pallickara, Shrideep Pallickara |
IEEE Big Data | 5 |
| 2020 | An Autocorrelation-based LSTM-Autoencoder for Anomaly Detection on Time-Series DataabstractData quality significantly impacts the results of data analytics. Researchers have proposed machine learning based anomaly detection techniques to identify incorrect data. Existing approaches fail to (1) identify the underlying domain constraints violated by the anomalous data, and (2) generate explanations of these violations in a form comprehensible to domain experts. We propose IDEAL, which is an LSTM-Autoencoder based approach that detects anomalies in multivariate time-series data, generates domain constraints, and reports subsequences that violate the constraints as anomalies. We propose an automated autocorrelation-based windowing approach to adjust the network input size, thereby improving the correctness and performance of constraint discovery over manual and brute-force approaches. The anomalies are visualized in a manner comprehensible to domain experts in the form of decision trees extracted from a random forest classifier. Domain experts can then provide feedback to retrain the learning model and improve the accuracy of the process. We evaluate the effectiveness of IDEAL using datasets from Yahoo servers, NASA Shuttle, and Colorado State University Energy Institute. We demonstrate that IDEAL can detect previously known anomalies from these datasets. Using mutation analysis, we show that IDEAL can detect different types of injected faults. We also demonstrate that the accuracy improves after incorporating domain expert feedback. Hajar Homayouni, Sudipto Ghosh 0001, Indrakshi Ray, Shlok Gondalia, Jerry Duggan, Michael G. Kahn |
IEEE BigData | 2 |
| 2019 | An Interactive Data Quality Test Approach for Constraint Discovery and Fault DetectionabstractData quality tests validate heterogeneous data to detect violations of syntactic and semantic constraints. The specification of these constraints can be incomplete because domain experts typically specify them in an ad hoc manner. Existing automated test approaches can generate false alarms and do not explain the constraint violations while reporting faulty data records. In previous work, we proposed ADQuaTe, which is an automated data quality test approach that uses an unsupervised deep learning techni que (1) to discover constraints from big datasets that may have been missed by experts, and (2) to label as suspicious those records that violate the constraints. These records are grouped and explanations for constraint violations are presented to domain experts who determine whether or not the groups are actually faulty. This paper presents ADQuaTe2, which extends ADQuaTe to use an interactive learning technique that incorporates expert feedback to retrain the learning model and improve the accuracy of constraint discovery and fault detection. We evaluate the effectiveness of the approach on real-world datasets from a health data warehouse and a plant diagnosis database. We also use datasets with known faults from the UCI repository to evaluate the improvement in the accuracy of the approach after incorporating ground truth knowledge. Hajar Homayouni, Sudipto Ghosh 0001, Indrakshi Ray, Michael G. Kahn |
IEEE BigData | 2 |
| 2018 | An Approach for Testing the Extract-Transform-Load Process in Data Warehouse SystemsabstractProQuest powers research in academic, corporate, government, public and school libraries around the world with unique content. Explore millions of resources from scholarly journals, books, newspapers, videos and more. Hajar Homayouni, Sudipto Ghosh 0001, Indrakshi Ray |
IDEAS | 2 |
| 2006 | Developing Distributed Services Using an Aspect Oriented Model Driven FrameworkabstractTo manage the development of cooperative information systems that support the dynamics and mobility of modern businesses, separation of concern mechanisms and abstractions are needed. Model driven development (MDD) approaches utilize abstraction and transformation to handle complexity. In MDD, specifying transformations between models at various levels of abstraction can be a complex task. Specifying transformations for pervasive system services that are tangled with other system services is particularly difficult because the elements to be transformed are distributed across a model. This paper presents an aspect oriented model driven framework (AOMDF) that facilitates separation of pervasive services and supports their transformation across different levels of abstraction. The framework facilitates composition of pervasive services with enterprise services at various levels of abstraction. The framework is illustrated using an example in which a platform independent model of a banking service is transformed to a platform specific model. Arnor Solberg, Devon M. Simmonds, Y. Raghu Reddy, Robert B. France, Sudipto Ghosh 0001, Jan Øyvind Aagedal |
Int. J. Cooperative Inf. Syst. | 5 |