EDBT 2026 Demo / reviewers in the wild / expert
Sanjay Ranka
dblp:r/SanjayRanka
· DBLP profile ↗
22ranked-venue papers in the field
0as first author
8since 2021 · last 2025
0000-0003-4886-1988ORCID · verified
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 13Database Systems & Data Management · 5Big Data, Cloud & Distributed Data Systems · 4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Machine Learning Techniques for Data Reduction of Climate Applications
Xiao Li 0048, Qian Gong, Jaemoon Lee, Scott Klasky, Anand Rangarajan 0001, Sanjay Ranka |
PAKDD (1) | 6 |
| 2025 | Foundation Model for Lossy Compression of Spatiotemporal Scientific Data
Xiao Li 0048, Jaemoon Lee, Anand Rangarajan 0001, Sanjay Ranka |
PAKDD (6) | 4 |
| 2025 | Evaluating Generative Vehicle Trajectory Models for Traffic Intersection Dynamics
Yash Ranjan, Rahul Sengupta, Anand Rangarajan 0001, Sanjay Ranka |
PAKDD (6) | 4 |
| 2024 | Attention Based Machine Learning Methods for Data Reduction with Guaranteed Error BoundsabstractScientific applications in fields such as high energy physics, computational fluid dynamics, and climate science generate vast amounts of data at high velocities. This exponential growth in data production is surpassing the advancements in computing power, network capabilities, and storage capacities. To address this challenge, data compression or reduction techniques are crucial. These scientific datasets have underlying data structures that consist of structured and block structured multidimensional meshes where each grid point corresponds to a tensor. It is important that data reduction techniques leverage strong spatial and temporal correlations that are ubiquitous in these applications. Additionally, applications such as CFD, process tensors comprising hundred plus species and their attributes at each grid point. Reduction techniques should be able to leverage interrelationships between the elements in each tensor.In this paper, we propose an attention-based hierarchical compression method utilizing a block-wise compression setup. We introduce an attention-based hyper-block autoencoder to capture inter-block correlations, followed by a block-wise encoder to capture block-specific information. A PCA-based post-processing step is employed to guarantee error bounds for each data block. Our method effectively captures both spatiotemporal and inter-variable correlations within and between data blocks. Compared to the state-of-the-art SZ3, our method achieves up to 8× higher compression ratio on the multi-variable S3D dataset. When evaluated on single-variable setups using the E3SM and XGC datasets, our method still achieves up to 3× and 2× higher compression ratio, respectively. Xiao Li 0048, Jaemoon Lee, Anand Rangarajan 0001, Sanjay Ranka |
IEEE Big Data | 4 |
| 2024 | Guaranteeing Error Bounds with Preservation of Derived Quantities in Compressive AutoencodersabstractScientists conduct large-scale simulations to compute derived quantities from primary data. Thus, it is crucial that data compression techniques maintain bounded errors on these derived quantities or quantities of interest (QoI). Despite the notable performance of recent learned image/video compression approaches using neural networks, they do not guarantee reconstruction errors and cannot manage QoI. This work introduces the Guaranteed Autoencoder with Preserved QoI (GAEQ), which utilizes the interpretation that neural networks with piecewise linear units (PLUs) can be interpreted as a set of linear operators [1] . Although the operators are instance-specific, many instances share the same operator if they fall into the same region of the tessellation formed by PLUs. Jaemoon Lee, Anand Rangarajan 0001, Sanjay Ranka |
DCC | 3 |
| 2024 | Hybrid Approaches for Data Reduction of Spatiotemporal Scientific ApplicationsabstractScientists conduct large-scale simulations to compute derived quantities from primary data. Thus, it is crucial that data compression techniques maintain bounded errors on these derived quantities or quantities of interest (QOI). For many spatiotemporal applications, these QOIs are binary in nature and represent presence or absence of a physical phenomenon. In this work, we propose to use a hybrid approah for differential compression for such applications. We use a neural network (NN) approach to determine regions-of-interest (ROIs) where the binary QOIs are going to be prevalent. This is then used with traditional approaches that compress at a lower level (and higher accuracy) for these ROIs as compared to other regions. Xiao Li 0048, Qian Gong, Jaemoon Lee, Scott Klasky, Anand Rangarajan 0001, Sanjay Ranka |
DCC | 6 |
| 2023 | Constrained Autoencoders: Incorporating equality constraints in learned scientific data compressionabstractIn scientific data compression, it is crucial to preserve Quantities of Interest (QoI) derived from the data for accurate post-analysis of scientific applications. In this work, we present Constrained Autoencoders (CAEs) where we impose linear QoI as constraints on neural network activations. We circumvent the difficulty of using standard convex optimization methods on the output predictor in the context of autoencoder-driven compression. Jaemoon Lee, Anand Rangarajan 0001, Sanjay Ranka |
DCC | 3 |
| 2022 | Region-adaptive, Error-controlled Scientific Data Compression using Multilevel DecompositionabstractThe increase of computer processing speed is significantly outpacing improvements in network and storage bandwidth, leading to the big data challenge in modern science, where scientific applications can quickly generate much more data than that can be transferred and stored. As a result, big scientific data must be reduced by a few orders of magnitude while the accuracy of the reduced data needs to be guaranteed for further scientific explorations. Moreover, scientists are often interested in some specific spatial/temporal regions in their data, where higher accuracy is required. The locations of the regions requiring high accuracy can sometimes be prescribed based on application knowledge, while other times they must be estimated based on general spatial/temporal variation. In this paper, we develop a novel multilevel approach which allows users to impose region-wise compression error bounds. Our method utilizes the byproduct of a multilevel compressor to detect regions where details are rich and we provide the theoretical underpinning for region-wise error control. With spatially varying precision preservation, our approach can achieve significantly higher compression ratios than single-error bounded compression approaches and control errors in the regions of interest. Qian Gong, Ben Whitney, Chengzhu Zhang, Xin Liang 0001, Anand Rangarajan 0001, Jieyang Chen, Lipeng Wan 0001, Paul Ullrich, Qing Liu 0002, Robert Jacob, Sanjay Ranka, Scott Klasky |
SSDBM | 11 |
| 2015 | Scalable Machine Learning Approaches for Neighborhood Classification Using Very High Resolution Remote Sensing ImageryabstractUrban neighborhood classification using very high resolution (VHR) remote sensing imagery is a challenging and {\em emerging} application. A semi-supervised learning approach for identifying neighborhoods is presented which employs superpixel tessellation representations of VHR imagery. The image representation utilizes homogeneous and irregularly shaped regions termed superpixels and derives novel features based on intensity histograms, geometry, corner and superpixel density and scale of tessellation. The semi-supervised learning approach uses a support vector machine (SVM) to obtain a preliminary classification which is then subsequently refined using graph Laplacian propagation. Several intermediate stages in the pipeline are presented to showcase the important features of this approach. We evaluated this approach on four different geographic settings with varying neighborhood types and compared it with the recent Gaussian Multiple Learning algorithm. This evaluation shows several advantages, including model building, accuracy, and efficiency which makes it a great choice for deployment in large scale applications like global human settlement mapping and population distribution (e.g., LandScan), and change detection. Manu Sethi, Yupeng Yan, Anand Rangarajan 0001, Ranga Raju Vatsavai, Sanjay Ranka |
KDD | 5 |
| 2010 | Mixture models for learning low-dimensional roles in high-dimensional dataabstractArchived data often describe entities that participate in multiple roles. Each of these roles may influence various aspects of the data. For example, a register transaction collected at a retail store may have been initiated by a person who is a woman, a mother, an avid reader, and an action movie fan. Each of these roles can influence various aspects of the customer's purchase: the fact that the customer is a mother may greatly influence the purchase of a toddler-sized pair of pants, but have no influence on the purchase of an action-adventure novel. The fact that the customer is an action move fan and an avid reader may influence the purchase of the novel, but will have no effect on the purchase of a shirt. Manas Somaiya, Chris Jermaine, Sanjay Ranka |
KDD | 3 |
| 2010 | A Model-Agnostic Framework for Fast Spatial Anomaly DetectionabstractGiven a spatial dataset placed on an n × n grid, our goal is to find the rectangular regions within which subsets of the dataset exhibit anomalous behavior. We develop algorithms that, given any user-supplied arbitrary likelihood function, conduct a likelihood ratio hypothesis test (LRT) over each rectangular region in the grid, rank all of the rectangles based on the computed LRT statistics, and return the top few most interesting rectangles. To speed this process, we develop methods to prune rectangles without computing their associated LRT statistics. Mingxi Wu, Chris Jermaine, Sanjay Ranka, Xiuyao Song, John Gums |
ACM Trans. Knowl. Discov. Data | 3 |
| 2009 | A LRT framework for fast spatial anomaly detectionabstractGiven a spatial data set placed on an n x n grid, our goal is to find the rectangular regions within which subsets of the data set exhibit anomalous behavior. We develop algorithms that, given any user-supplied arbitrary likelihood function, conduct a likelihood ratio hypothesis test (LRT) over each rectangular region in the grid, rank all of the rectangles based on the computed LRT statistics, and return the top few most interesting rectangles. To speed this process, we develop methods to prune rectangles without computing their associated LRT statistics. Mingxi Wu, Xiuyao Song, Chris Jermaine, Sanjay Ranka, John Gums |
KDD | 4 |
| 2008 | A bayesian mixture model with linear regression mixing proportionsabstractClassic mixture models assume that the prevalence of the various mixture components is fixed and does not vary over time. This presents problems for applications where the goal is to learn how complex data distributions evolve. We develop models and Bayesian learning algorithms for inferring the temporal trends of the components in a mixture model as a function of time. We show the utility of our models by applying them to the real-life problem of tracking changes in the rates of antibiotic resistance in Escherichia coli and Staphylococcus aureus. The results show that our methods can derive meaningful temporal antibiotic resistance patterns. Xiuyao Song, Chris Jermaine, Sanjay Ranka, John Gums |
KDD | 3 |
| 2008 | Learning correlations using the mixture-of-subsets modelabstractUsing a mixture of random variables to model data is a tried-and-tested method common in data mining, machine learning, and statistics. By using mixture modeling it is often possible to accurately model even complex, multimodal data via very simple components. However, the classical mixture model assumes that a data point is generated by a single component in the model. A lot of datasets can be modeled closer to the underlying reality if we drop this restriction. We propose a probabilistic framework, the mixture-of-subsets (MOS) model , by making two fundamental changes to the classical mixture model. First, we allow a data point to be generated by a set of components, rather than just a single component. Next, we limit the number of data attributes that each component can influence. We also propose an EM framework to learn the MOS model from a dataset, and experimentally evaluate it on real, high-dimensional datasets. Our results show that the MOS model learned from the data represents the underlying nature of the data accurately. Manas Somaiya, Chris Jermaine, Sanjay Ranka |
ACM Trans. Knowl. Discov. Data | 3 |
| 2007 | Statistical change detection for multi-dimensional dataabstractThis paper deals with detecting change of distribution in multi-dimensional data sets. We use sequential hypothesis testing methods from statistics to define a general, Monte-Carlo framework for solving this problem. We also define a specific statistical test for distributional change within the framework, that we call the density test. Our experimental results show that the density test has substantially more power than the two existing methods for multi-dimensional change detection. Xiuyao Song, Mingxi Wu, Chris Jermaine, Sanjay Ranka |
KDD | 4 |
| 2007 | Conditional Anomaly DetectionabstractWhen anomaly detection software is used as a data analysis tool, finding the hardest-to-detect anomalies is not the most critical task. Rather, it is often more important to make sure that those anomalies that are reported to the user are in fact interesting. If too many unremarkable data points are returned to the user labeled as candidate anomalies, the software can soon fall into disuse. One way to ensure that returned anomalies are useful is to make use of domain knowledge provided by the user. Often, the data in question includes a set of environmental attributes whose values a user would never consider to be directly indicative of an anomaly. However, such attributes cannot be ignored because they have a direct effect on the expected distribution of the result attributes whose values can indicate an anomalous observation. This paper describes a general purpose method called conditional anomaly detection for taking such differences among attributes into account, and proposes three different expectation-maximization algorithms for learning the model that is used in conditional anomaly detection. Experiments with more than 13 different data sets compare our algorithms with several other more standard methods for outlier or anomaly detection Xiuyao Song, Mingxi Wu, Chris Jermaine, Sanjay Ranka |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2000 | Performance Analysis of Parallel Query Processing Algorithms for Object-Oriented DatabasesabstractTwo types of parallel processing and optimization algorithms for processing object-oriented databases are the hybrid-hash pointer-based (HHP) algorithms and multi-wavefront (MWF) algorithms. We analyze these two algorithms and develop analytical formulas to capture their main performance features. We study their performance in three application environments, characterized by large databases having many object classes, each of which, respectively, (1) contains a large number of instances; (2) contains a relatively small number of instances; and (3) is of varying size. A horizontal data partitioning strategy is used in (1). A class-per-node assignment strategy is used in (2). In (3), object classes are partitioned horizontally and assigned to a varying number of processors depending on their different sizes. The MWF algorithm has three distinguishing features which contribute to its better performance: (a) a two-phase processing strategy, (b) vertical partitioning of horizontal segments, and (c) dynamic determination of the collision point in MWF propagations, which results in an optimized query execution plan. If these features are adopted by an HHP algorithm, its performance is comparable with that of the MWF algorithm because the difference in CPU time between them is negligible. The computing environment is a network of workstations having a shared-nothing architecture. The schema and some queries selected from the OO7 benchmark are used in the performance analyses and comparisons. The queries are modified slightly in different data environments in order to reflect the features of diverse database applications. Stanley Y. W. Su, Sanjay Ranka |
IEEE Trans. Knowl. Data Eng. | 2 |
| 1999 | An Efficient Space-Partitioning Based Algorithm for the K-Means Clustering
Khaled Alsabti, Sanjay Ranka |
PAKDD | 2 |
| 1998 | CLOUDS: A Decision Tree Classifier for Large Datasets
Khaled Alsabti, Sanjay Ranka |
KDD | 2 |
| 1997 | An Efficient Algorithm for the Incremental Updation of Association Rules in Large Databases
Shiby Thomas, Sreenath Bodagala, Khaled Alsabti, Sanjay Ranka |
KDD | 4 |
| 1997 | A One-Pass Algorithm for Accurately Estimating Quantiles for Disk-Resident Data
Khaled Alsabti, Sanjay Ranka |
VLDB | 2 |
| 1994 | A Space-and-Time-Efficient Codeing Algorithm for Lattice ComputationsabstractWe present an encoding algorithm for lattices that significantly reduces space requirements while allowing fast computations of least upper bounds and greatest lower bounds of pairs of elements. We analyze the algorithms for encoding, LUB and GLB computations, and prove their correctness. Empirical experiments reveal that our method is significantly more space efficient than the transitive closure method, and the saving becomes increasingly more important as the size of the lattice increases.> Deb Dutta Ganguly, Chilukuri K. Mohan, Sanjay Ranka |
IEEE Trans. Knowl. Data Eng. | 3 |