VLDB 2026 Research / reviewers in the wild / expert
Paahuni Khandelwal
dblp:252/7188
· DBLP profile ↗
10ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0002-9026-5242ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 3 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Science-Informed Multitask Transformer for Soil Property Prediction from FTIR SpectroscopyabstractMonitoring soil health at scale requires tools that are both scientifically robust and computationally efficient. Accurate and efficient estimation of soil chemical properties is essential for sustainable agricultural practices and environmental management. Traditional laboratory-based soil testing methods are labor-intensive, time-consuming, and often impractical for large-scale analysis. Although mid-infrared (MIR) spectroscopy offers a rapid and cost-effective alternative, conventional modeling techniques suffer from inconsistent predictive performance across diverse soil properties. To address this gap, we introduce FTIRNet, a novel multi-task Transformer-based architecture designed to predict multiple soil chemical properties from Fourier-transform infrared (FTIR) spectral data. FTIRNet employs a shared encoder block with task-specific Transformer encoders, enabling the model to learn both general and task-focused spectral representations. These multi-task Transformer encoders incorporate science-informed task relation learning through specialized fusion layers that capture known biogeochemical interdependencies among soil properties. Across a broad set of soil properties, FTIRNet demonstrates consistently superior predictive accuracy compared to traditional approaches. Andrei Bachinin, Rupasree Dey, Paahuni Khandelwal, Sam Leuthold, M. Francesca Cotrufo, Shrideep Pallickara, Sangmi Lee Pallickara |
eScience | 3 |
| 2024 | Periscope: A Framework for Visualizations of Multiresolution Spatiotemporal Data at ScaleabstractThe crux of this study is to support browser-based visualizations of spatiotemporally evolving phenomena. Such phenomena arise in myriad domains spanning terrestrial, oceanic, and atmospheric processes. The data are voluminous, have diverse representational formats and projection systems, and are multivariate. We rely on a novel mix of tiling, caching, compression, perceptual limits, speculative prefetching, and dynamic generation of tiles. Our refinements at the client and server-side work in concert with each other to leverage client-side resources, minimize duplicate processing, and effective prefetching to ensure interactive explorations at scale. Our benchmarks profiled several aspects of our methodology and demonstrate the suitability of our refinements. Everett Lewark, Matthew Young, Paahuni Khandelwal, Sangmi Lee Pallickara, Shrideep Pallickara |
IEEE Big Data | 3 |
| 2024 | DeepSoil: A Science-guided Framework for Generating High Precision Soil Moisture Maps by Reconciling Measurement Profiles Across In-situ and Remote Sensing DataabstractSoil moisture plays a critical role in several domains and can be used to inform decision-making in agricultural settings, drought forecasting, forest fire predictions, and water conservation. Soil moisture is measured using in-situ and remote-sensing equipment. Depending on the type of equipment that is used, some challenges must be reconciled, including the density of observations, the measurement precision, and the resolutions at which these measurements are available. In particular, in-situ measurements are high-precision but sparse, while remote sensing measurements benefit from spatial coverage, albeit at lower precision and coarser resolutions. The crux of this study is to produce higher-precision soil moisture estimates at high resolutions (30m). Our methodology combines scientific models, deep networks, topographical characteristics, and information about ambient conditions alongside both in-situ and remote sensing data to accomplish this. Domain science infuses several aspects of our methodology. Our empirical benchmarks profile several aspects and demonstrate that our methodology accounts for spatial variability while accounting for both static (soil properties and elevation) and dynamically varying phenomena to generate accurate, high-precision 30m resolution soil moisture content maps. Paahuni Khandelwal, Sangmi Lee Pallickara, Shrideep Pallickara |
SIGSPATIAL/GIS | 1 |
| 2023 | ARGUS: Rapid Wildfire Tracking Using Satellite Data CollectionsabstractInteractive visual analytics over distributed systems housing voluminous datasets is hindered by three main factors - disk and network I/O, and data processing overhead. Requests over geospatial data are prone to erratic query load and hotspots due to users' simultaneous interest over a small sub-domain of the overall data space at a time. Interactive analytics in a distributed setting is further hindered in cases of voluminous datasets with large/high-dimensional data objects, such as multi-spectral satellite imagery. The size of the data objects prohibits efficient caching mechanisms that could significantly reduce response latencies. Additionally, extracting information from these large data objects incurs significant data processing overheads and they often entail resource-intensive computational methods. Here, we present our framework, Argus,that extracts low-dimensional representation (embeddings) of high-dimensional satellite images during ingestion and houses them in the cache for use in model-driven analysis relating to wildfire detection. These embeddings are versatile and are used to perform model-based extraction of analytical information for a set of different scenarios, to reduce the high computational costs that are involved with typical transformations over high-dimensional datasets. The models for each such analytical process are trained in a distributed manner in a connected, multi-task learning fashion, along with the encoder network that generates the original embeddings. Saptashwa Mitra, Paahuni Khandelwal, Shrideep Pallickara, Sangmi Lee Pallickara |
CLOUD | 2 |
| 2023 | DISCERN: Leveraging Knowledge Distillation to Generate High Resolution Soil Moisture Estimation from Coarse Satellite DataabstractAccurate estimation of soil moisture is crucial for efficient agricultural management and environmental monitoring. However, the task of predicting soil moisture levels becomes challenging in regions with limited data availability. In this study, we propose a knowledge distillation-based deep learning approach to enhance soil moisture prediction with machine learning apporach using the low resolution but wide coverage soil moisture Active Passive (SMAP) satellite data.Our framework leverages the knowledge distillation, where a high-capacity teacehr model (VGG13) which is pre-traineed on a large dataset (SMAP) and a lightweight student model (ResNet8) which is then trained on sensor-based highly accurate but extremely sparse station data. The student model benefits from the distilled knowledge of the teacher model, acquiring a deeper understanding of the underlying patterns and relationships in the data.The space-efficient student model significantly reduces the inference time with high prediction accuracy and demonstrates the potential benefit to agricultural management, water resource planning, and ecological studies by providing accurate and reliable soil moisture predictions in data-scarce regions. Our findings reveal how to identify performant settings for achieving the best trade-off between accuracy and model complexity. Abdul Matin, Paahuni Khandelwal, Shrideep Pallickara, Sangmi Lee Pallickara |
IEEE Big Data | 2 |
| 2023 | Enabling Fast, Effective Visualization of Voluminous Gridded Spatial DatasetsabstractGridded spatial datasets arise naturally in environmental, climatic, meteorological, and ecological settings. Each grid point encapsulates a vector of variables representing different measures of interest. Gridded datasets tend to be voluminous since they encapsulate observations for long timescales. Visualizing such datasets poses significant challenges stemming from the need to preserve interactivity, manage I/O overheads, and cope with data volumes. Here we present our methodology to significantly alleviate I/O requirements by leveraging deep neural network-based models and a distributed, in-memory cache to facilitate interactive visualizations. Our benchmarks demonstrate that deploying our lightweight models coupled with back-end caching and prefetching schemes can reduce the client's query response time by 92.3% while maintaining a high perceptual quality with a PSNR (peak signal-to-noise ratio) of 38.7 dB. Paahuni Khandelwal, Menuka Warushavithana, Sangmi Lee Pallickara, Shrideep Pallickara |
CCGrid | 1 |
| 2022 | CloudNet: A Deep Learning Approach for Mitigating Occlusions in Landsat-8 Imagery using Data CoalescenceabstractMulti-spectral satellite images that remotely sense the Earth's surface at regular intervals are often contaminated due to occlusion by clouds. Remote sensing imagery captured via satellites, drones, and aircraft has successfully influenced a wide range of fields such as monitoring vegetation health, tracking droughts, and weather forecasting, among others. Researchers studying the Earth's surface are often hindered while gathering reliable observations due to contaminated reflectance values that are sensitive to thin, thick, and cirrus clouds, as well as their shadows. In this study, we propose a deep learning network architecture, CloudNet, to alleviate cloud-occluded remote sensing imagery captured by Landsat-8 satellite for both visible and non-visible spectral bands. We propose a deep neural network model trained on a distributed storage cluster that leverages historical trends within Landsat-8 imagery while complementing this analysis with high-resolution Sentinel-2 imagery. Our empirical benchmarks profile the efficiency of the CloudNet model with a range of cloud-occluded pixels in the input image. We further compare our CloudNet's performance with state-of-the-art deep learning approaches such as SpAGAN and Resnet. We propose a novel method, dynamic hierarchical transfer learning, to reduce computational resource requirements while training the model to achieve the desired accuracy. Our model regenerates features of cloudy images with a high PSNR accuracy of 34.28 dB. Paahuni Khandelwal, Samuel Armstrong, Abdul Matin, Shrideep Pallickara, Sangmi Lee Pallickara |
e-Science | 1 |
| 2021 | Mind the Gap: Generating Imputations for Satellite Data Collections at Myriad Spatiotemporal ScopesabstractHyperspectral satellite data collections have been successfully leveraged in many domains such as meteorology, agriculture, forestry, and disaster management. There is also a collection of publicly available satellite observation networks. However, gaps in scanning frequencies and inadequate spatial resolutions limit the capabilities of geoscience applications. In this study, we target the temporal sparsity of high-resolution satellite images. In particular, we propose a novel methodology to estimate high-resolution images between scheduled scans. Our model SATnet, falls broadly within the class of Generative Adversarial Networks. SATnet allows us to generate accurate high-resolution, high-frequency satellite data at diverse spatial extents. SATnet achieves this by learning relations between a sequence of high-resolution/low-frequency satellite imageries (from Sentinel-2) and an ancillary satellite image that is high-frequency/low-resolution (from MODIS). Our benchmarks demonstrate that SATnet outperforms existing approaches such as ConvLSTMs, Dynamic Filter Network, and TrajGRU with a PSNR accuracy of 31.82. We trained and deployed SATnet over a distributed storage cluster to support the high-throughput generation of imputed satellite imagery via query evaluations. Our methodology preserves geospatial proximity and facilitates the dynamic construction of satellite imagery at a particular timestamp for arbitrary spatial scopes. Paahuni Khandelwal, Daniel Rammer, Shrideep Pallickara, Sangmi Lee Pallickara |
CCGRID | 1 |
| 2020 | Lightweight, Embeddings Based Storage and Model Construction Over Satellite Data CollectionsabstractThere has been a substantial growth in remotely sensed hyperspectral satellite imagery. These data offer opportunities to understand phenomena and inform decision making. The nature of these collections introduces challenges stemming from their volumes, variety, and spatiotemporal resolutions. The crux of this study is to facilitate effective training of deep learning models over satellite data collections. We describe our novel embeddings (multidimensional latent space representations) based approach to effectively support model training, refinement, and inferences. We rigorously explore several aspects relating to embeddings, including their dimensionality, single vs multiple bands, and preservation of inter-band metrics. We also incorporate support for transfer learning over spatiotemporal scopes to address issues relating to cold start and alleviate resource pressure. Our methodology addresses disk, network, CPU/GPU, and accuracy implications of several aspects relating to model construction. Our empirical benchmarks assess the suitability of our methodology using the MODIS and Sentinel-2 satellite data. We demonstrate that our methodology reduces storage requirements by more than 10,000x and reduces model construction times by 75%. Kevin Bruhwiler, Paahuni Khandelwal, Daniel Rammer, Samuel Armstrong, Sangmi Lee Pallickara, Shrideep Pallickara |
IEEE BigData | 2 |
| 2019 | STASH : Fast Hierarchical Aggregation Queries for Effective Visual Spatiotemporal ExplorationsabstractThe proliferation of sensors and observational instruments enable scientists to explore natural, spatiotemporal phenomena via explorative analysis and advanced modeling. Geospatial visualization, in particular, is an intuitive tool to identify patterns, enhance understanding of the data, and plan for subsequent analysis. However, seamless interactions between end-user devices and the sheer volume of data have been a challenge due to the limited bandwidth and data access latencies.In this paper, we introduce Stash, a distributed, in-memory cache for hierarchical aggregation and query evaluations. Stash is a middleware which can be loaded on top of a distributed file system. Users perform queries from a lightweight visualization interface at the front-end and the evaluations occur over the back-end storage system housing the raw data over which summarization and subsequent visualizations are to be performed. Stash facilitates fast exploratory analytics by caching relevant past query results based on their frequency and freshness to assist similar, future queries and avoid expensive disk I/O and network usage, thus reducing their latency. Additionally, Stash handles any hotspot that might result from a spike in user requests due to the spatial and temporal locality of their access patterns.Our empirical benchmarks show that a Stash-enabled system reduces query latency of a basic system by over 5-folds and brings it down to interactive speed even for large country-sized spatiotemporal queries. We have contrasted Stash with existing cache-enabled analytics engines, such as ElasticSearch, and found that our STASH-enabled system reduced the aggregation query latency up to ~70%. STASH also alleviated skewed workloads through its dynamic replication scheme and improved throughput by ~40% in hotspot scenarios. Saptashwa Mitra, Paahuni Khandelwal, Shrideep Pallickara, Sangmi Lee Pallickara |
CLUSTER | 2 |