VLDB 2026 Research / reviewers in the wild / expert
Anand Rangarajan 0001
dblp:90/6511-1
· DBLP profile ↗
10ranked-venue papers in the field
0as first author
8since 2021 · last 2025
0000-0001-8695-8436ORCID · verified
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 5Big Data, Cloud & Distributed Data Systems · 4Database Systems & Data Management · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Machine Learning Techniques for Data Reduction of Climate Applications
Xiao Li 0048, Qian Gong, Jaemoon Lee, Scott Klasky, Anand Rangarajan 0001, Sanjay Ranka |
PAKDD (1) | 5 |
| 2025 | Foundation Model for Lossy Compression of Spatiotemporal Scientific Data
Xiao Li 0048, Jaemoon Lee, Anand Rangarajan 0001, Sanjay Ranka |
PAKDD (6) | 3 |
| 2025 | Evaluating Generative Vehicle Trajectory Models for Traffic Intersection Dynamics
Yash Ranjan, Rahul Sengupta, Anand Rangarajan 0001, Sanjay Ranka |
PAKDD (6) | 3 |
| 2024 | Attention Based Machine Learning Methods for Data Reduction with Guaranteed Error BoundsabstractScientific applications in fields such as high energy physics, computational fluid dynamics, and climate science generate vast amounts of data at high velocities. This exponential growth in data production is surpassing the advancements in computing power, network capabilities, and storage capacities. To address this challenge, data compression or reduction techniques are crucial. These scientific datasets have underlying data structures that consist of structured and block structured multidimensional meshes where each grid point corresponds to a tensor. It is important that data reduction techniques leverage strong spatial and temporal correlations that are ubiquitous in these applications. Additionally, applications such as CFD, process tensors comprising hundred plus species and their attributes at each grid point. Reduction techniques should be able to leverage interrelationships between the elements in each tensor.In this paper, we propose an attention-based hierarchical compression method utilizing a block-wise compression setup. We introduce an attention-based hyper-block autoencoder to capture inter-block correlations, followed by a block-wise encoder to capture block-specific information. A PCA-based post-processing step is employed to guarantee error bounds for each data block. Our method effectively captures both spatiotemporal and inter-variable correlations within and between data blocks. Compared to the state-of-the-art SZ3, our method achieves up to 8× higher compression ratio on the multi-variable S3D dataset. When evaluated on single-variable setups using the E3SM and XGC datasets, our method still achieves up to 3× and 2× higher compression ratio, respectively. Xiao Li 0048, Jaemoon Lee, Anand Rangarajan 0001, Sanjay Ranka |
IEEE Big Data | 3 |
| 2024 | Guaranteeing Error Bounds with Preservation of Derived Quantities in Compressive AutoencodersabstractScientists conduct large-scale simulations to compute derived quantities from primary data. Thus, it is crucial that data compression techniques maintain bounded errors on these derived quantities or quantities of interest (QoI). Despite the notable performance of recent learned image/video compression approaches using neural networks, they do not guarantee reconstruction errors and cannot manage QoI. This work introduces the Guaranteed Autoencoder with Preserved QoI (GAEQ), which utilizes the interpretation that neural networks with piecewise linear units (PLUs) can be interpreted as a set of linear operators [1] . Although the operators are instance-specific, many instances share the same operator if they fall into the same region of the tessellation formed by PLUs. Jaemoon Lee, Anand Rangarajan 0001, Sanjay Ranka |
DCC | 2 |
| 2024 | Hybrid Approaches for Data Reduction of Spatiotemporal Scientific ApplicationsabstractScientists conduct large-scale simulations to compute derived quantities from primary data. Thus, it is crucial that data compression techniques maintain bounded errors on these derived quantities or quantities of interest (QOI). For many spatiotemporal applications, these QOIs are binary in nature and represent presence or absence of a physical phenomenon. In this work, we propose to use a hybrid approah for differential compression for such applications. We use a neural network (NN) approach to determine regions-of-interest (ROIs) where the binary QOIs are going to be prevalent. This is then used with traditional approaches that compress at a lower level (and higher accuracy) for these ROIs as compared to other regions. Xiao Li 0048, Qian Gong, Jaemoon Lee, Scott Klasky, Anand Rangarajan 0001, Sanjay Ranka |
DCC | 5 |
| 2023 | Constrained Autoencoders: Incorporating equality constraints in learned scientific data compressionabstractIn scientific data compression, it is crucial to preserve Quantities of Interest (QoI) derived from the data for accurate post-analysis of scientific applications. In this work, we present Constrained Autoencoders (CAEs) where we impose linear QoI as constraints on neural network activations. We circumvent the difficulty of using standard convex optimization methods on the output predictor in the context of autoencoder-driven compression. Jaemoon Lee, Anand Rangarajan 0001, Sanjay Ranka |
DCC | 2 |
| 2022 | Region-adaptive, Error-controlled Scientific Data Compression using Multilevel DecompositionabstractThe increase of computer processing speed is significantly outpacing improvements in network and storage bandwidth, leading to the big data challenge in modern science, where scientific applications can quickly generate much more data than that can be transferred and stored. As a result, big scientific data must be reduced by a few orders of magnitude while the accuracy of the reduced data needs to be guaranteed for further scientific explorations. Moreover, scientists are often interested in some specific spatial/temporal regions in their data, where higher accuracy is required. The locations of the regions requiring high accuracy can sometimes be prescribed based on application knowledge, while other times they must be estimated based on general spatial/temporal variation. In this paper, we develop a novel multilevel approach which allows users to impose region-wise compression error bounds. Our method utilizes the byproduct of a multilevel compressor to detect regions where details are rich and we provide the theoretical underpinning for region-wise error control. With spatially varying precision preservation, our approach can achieve significantly higher compression ratios than single-error bounded compression approaches and control errors in the regions of interest. Qian Gong, Ben Whitney, Chengzhu Zhang, Xin Liang 0001, Anand Rangarajan 0001, Jieyang Chen, Lipeng Wan 0001, Paul Ullrich, Qing Liu 0002, Robert Jacob, Sanjay Ranka, Scott Klasky |
SSDBM | 5 |
| 2015 | Scalable Machine Learning Approaches for Neighborhood Classification Using Very High Resolution Remote Sensing ImageryabstractUrban neighborhood classification using very high resolution (VHR) remote sensing imagery is a challenging and {\em emerging} application. A semi-supervised learning approach for identifying neighborhoods is presented which employs superpixel tessellation representations of VHR imagery. The image representation utilizes homogeneous and irregularly shaped regions termed superpixels and derives novel features based on intensity histograms, geometry, corner and superpixel density and scale of tessellation. The semi-supervised learning approach uses a support vector machine (SVM) to obtain a preliminary classification which is then subsequently refined using graph Laplacian propagation. Several intermediate stages in the pipeline are presented to showcase the important features of this approach. We evaluated this approach on four different geographic settings with varying neighborhood types and compared it with the recent Gaussian Multiple Learning algorithm. This evaluation shows several advantages, including model building, accuracy, and efficiency which makes it a great choice for deployment in large scale applications like global human settlement mapping and population distribution (e.g., LandScan), and change detection. Manu Sethi, Yupeng Yan, Anand Rangarajan 0001, Ranga Raju Vatsavai, Sanjay Ranka |
KDD | 3 |
| 2015 | Discriminative Interpolation for Classification of Functional Data
Rana Haber, Anand Rangarajan 0001, Adrian M. Peter |
ECML/PKDD (1) | 2 |