EDBT 2026 Demo / reviewers in the wild / expert
Vandana Pursnani Janeja
dblp:44/111 · also Vandana P. Janeja
· DBLP profile ↗
30ranked-venue papers in the field
5as first author
11since 2021 · last 2025
0000-0003-0130-6135ORCID · verified
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 10 (1 first)Data Mining & Knowledge Discovery · 10 (3 first)Big Data, Cloud & Distributed Data Systems · 8 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 1Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improving Greenland Bed Topography Mapping with Uncertainty-Aware Graph Learning on Sparse Radar Data
Bayu Adhi Tama, Homayra Alam, Mostafa Cham, Omar Faruque, Jianwu Wang 0001, Vandana Pursnani Janeja |
IEEE Big Data | 6 |
| 2025 | Uncertainty-Aware Anomaly Detection in Spatiotemporal Climate Data [Experiment]abstractAccurate detection and quantification of uncertainty in anomalous climate events are needed to improve our understanding of climate extremes, particularly snow and ice melt processes in the polar regions. Despite the advances in anomaly detection methods, existing frameworks often neglect uncertainties inherent in spatiotemporal processes, leading to unreliable anomaly detection. We propose an uncertainty-aware anomaly detection framework that integrates measurement uncertainty and modeling bias. Our approach leverages the Three-Cornered-Hat (3CH) error variance estimator to quantify input uncertainties and incorporate them into the anomaly detection process through an uncertainty-weighted loss function and Monte Carlo Dropout (MCD) for total predictive uncertainty estimation. Using our approach, we detect and evaluate the uncertainty associated with anomalies from three surface melt products (ERA5, MAR, GEMB) of the Greenland Ice Sheet surface. Experiments on synthetic datasets and modeling output demonstrate that our uncertainty-aware method significantly enhances anomaly detection reliability, reducing false positives and improving detection confidence. Our results across the three models provide robust insights into the simulation of ice sheet surface melt dynamics, highlighting that GEMB, which includes the most complex physical representation of snow evolution, exhibits the strongest reliability in detecting melt regions with low uncertainty. Our findings provide insights into the simulation of ice sheet surface melt dynamics and underscore the importance of uncertainty quantification in Earth system modeling. Tolulope Ale, Ratnaksha Lele, Nicole-Jeanne Schlegel, Vandana Pursnani Janeja |
SIGSPATIAL/GIS | 4 |
| 2025 | Modeling Heterogeneity across Varying Spatial Extents: Discovering Linkages between Sea Ice Retreat and Ice Shelf Melt in the AntarcticabstractSpatial phenomena often exhibit heterogeneity across spatial extents and even in proximity. This spatial variation is complex to model, especially for large spatial extents that may vary, for example, ice shelves and sea ice. In this paper, we address this gap and, in particular, highlight its use in understanding linkages between sea ice retreat and Antarctic ice shelf (AIS) melt in the Antarctic. The Antarctic is losing ice at an unprecedented rate. While the role of atmospheric forcing and basal melting on both sea ice retreat and continental ice mass loss has been widely studied, how the retreat of sea ice affects the AIS mass loss has not been investigated. In fact, the link between the two has not been well established yet. Traditional models often treat sea ice and AIS as independent systems, limiting their ability to capture localized linkages and cascading feedbacks. To address this, we propose Spatial-Link, a novel graph-based modeling framework that quantifies spatial heterogeneity across varying spatial extents to capture the linkages between sea ice retreat and AIS melt. Our analysis shows how sea ice retreat evolves over an oceanic grid and gradually progresses to the ice shelf - establishing a direct linkage. To the best of our knowledge, this is the first-ever direct linkage methodology between sea ice retreat and AIS melt events. By integrating dynamic cryospheric processes into a unified framework, Spatial-Link offers a scalable, data-driven tool for improving the accuracy of sea-level rise projections and informing targeted climate adaptation strategies. Maloy Kumar Devnath, Sudip Chakraborty, Vandana Pursnani Janeja |
SIGSPATIAL/GIS | 3 |
| 2025 | Physics-Guided Multi-Contextual Learning: Understanding the Surface and Subsurface Processes in Southeast GreenlandabstractGreenland ice loss contributes approximately 0.7 mm per year to current sea level rise, making process attribution critical for future projections. Mass balance attribution in ice sheets requires distinguishing between surface processes (accumulation, ablation) and subsurface processes (submarine melting, dynamic discharge). Existing approaches lack systematic integration of glaciological process knowledge for robust spatial attribution. We present a physics-guided multi-contextual analysis framework. We construct process-specific variables using fundamental ice sheet mass balance principles. We create spatial neighborhoods through Voronoi polygon construction and feature similarity assessment, then apply Local Indicators of Spatial Association (LISA) to identify process dominance patterns. We advance beyond conventional spatial statistics by systematically deriving physics-informed indicators that isolate distinct mass balance components. We test our framework on Southeast Greenland using 18 years of reanalysis and satellite data (2004–2021). Our results show that subsurface processes control ice loss across 37–46% of the study area, concentrated in northern regions of Southeast Greenland, while surface processes dominate only 6–7% of the area in southern locations of Southeast Greenland. When we compare our spatial findings with the documented glacier behavior from recent studies in Greenland, we find strong agreement that validates our process attribution approach. Chhaya Kulkarni, Nicole-Jeanne Schlegel, Vandana Pursnani Janeja |
SIGSPATIAL/GIS | 3 |
| 2025 | DeepTopoNet: A Framework for Subglacial Topography Estimation on the Greenland Ice SheetsabstractMapping Greenland's subglacial topography is critical for projecting the future mass loss of the ice sheet and its contribution to global sea-level rise. However, the complex and sparse nature of observational data, particularly information about the bed topography under the ice sheet, significantly increases the uncertainty in model projections. Bed topography is traditionally measured by airborne ice-penetrating radars that measure the ice thickness directly underneath the aircraft, leaving data gaps of tens of kilometers in between flight lines. This study introduces a deep learning framework, DeepTopoNet, that integrates radar-derived ice thickness observations and BedMachine Greenland data through a novel dynamic loss-balancing mechanism. Among all efforts to reconstruct bed topography, BedMachine has emerged as one of the most widely used datasets, combining mass conservation principles and ice thickness measurements to generate high-resolution bed elevation estimates. The proposed loss function adaptively adjusts the weighting between radar and BedMachine's bed, ensuring robustness in areas with limited radar coverage while leveraging the high spatial resolution of BedMachine's bed estimates. Our approach incorporates gradient-based and trend surface features to enhance model performance and utilizes a convolutional neural network (CNN) architecture (i.e., BedTopoCNN) designed for subgrid-scale predictions. By systematically testing on the Upernavik Isstrøm) region in West Greenland, the model achieves high accuracy (MAE: 12.49 m, RMSE: 19.38 m, and R2: 0.99), outperforming baseline methods in reconstructing subglacial terrain. This work demonstrates the potential of deep learning in bridging observational gaps, providing a scalable and efficient solution to inferring subglacial topography. This framework paves the way for improved predictions of ice sheet flow and sea level rise. Bayu Adhi Tama, Mansa Krishna, Homayra Alam, Mostafa Cham, Omar Faruque, Gong Cheng 0004, Jianwu Wang 0001, Mathieu Morlighem, Vandana Pursnani Janeja |
SIGSPATIAL/GIS | 9 |
| 2025 | Advancing Climate Model Interpretability: Feature Attribution for Arctic Melt AnomaliesabstractThe focus of our work is improving the interpretability of anomalies in climate models and advancing our understanding of Arctic melt dynamics. The Arctic and Antarctic ice sheets are experiencing rapid surface melting and increased freshwater runoff, contributing significantly to global sea level rise. Understanding the mechanisms driving snowmelt in these regions is crucial. ERA5, a widely used reanalysis dataset in polar climate studies, offers extensive climate variables and global data assimilation. However, its snowmelt model employs an energy imbalance approach that may oversimplify the complexity of surface melt. In contrast, the Glacier Energy and Mass Balance (GEMB) model incorporates additional physical processes, such as snow accumulation, firn densification, and meltwater percolation/refreezing, providing a more detailed representation of surface melt dynamics. In this research, we focus on analyzing surface snowmelt dynamics of the Greenland Ice Sheet using feature attribution for anomalous melt events in ERA5 and GEMB models. We present a novel unsupervised attribution method leveraging counterfactual explanation method to analyze detected anomalies in ERA5 and GEMB. Our anomaly detection results are validated using MEaSUREs ground-truth data, and the attributions are evaluated against established and new feature ranking methods, including XGBoost, Shapley values, Random Forest, and Layer-wise Relevance Propagation. Our attribution framework identifies the physics behind each model and the climate features driving melt anomalies. These findings demonstrate the utility of our attribution method in enhancing the interpretability of anomalies in climate models and advancing our understanding of Arctic melt dynamics. Tolulope Ale, Nicole-Jeanne Schlegel, Vandana Pursnani Janeja |
ICDM | 3 |
| 2025 | Blue Sky: Expert-in-the-Loop Representation Learning Framework for Audio Anti-Spoofing: Multimodal, Multilingual, Multi-speaker, Multi-attack (4M) ScenariosabstractAudio spoofing has surged with the rise of generative artificial intelligence, posing a serious threat to online communication. Recent studies have shown promising avenues in detecting spoofed audio specifically those that use human expert knowledge in representation learning, but more work is needed to evaluate performance across various realistic scenarios that tend to pose challenges in spoofed audio detection. In this paper, we introduce a comprehensive framework for expert-in-the-loop representation learning for audio anti-spoofing that is robust enough to address four specific challenging scenarios. Multimodal, Multilingual, Multi-speaker, and Multi-attack (4M). Preliminary results demonstrate the framework’s potential effectiveness in audio anti-spoofing. Zahra Khanjani, Vandana Pursnani Janeja, Christine Mallinson, Sanjay Purushotham |
SDM | 2 |
| 2024 | Hybrid Ensemble Deep Graph Temporal Clustering for Spatiotemporal DataabstractThe increasing complexity of multidimensional spatiotemporal data presents significant challenges for clustering techniques, particularly in capturing intricate temporal, spatial, and heterogeneous patterns. This paper proposes a novel Hybrid Ensemble Deep Graph Temporal Clustering (HEDGTC) algorithm that integrates homogeneous and heterogeneous ensemble clustering models, leveraging both traditional and deep learning-based clustering approaches. The algorithm utilizes graph neural networks (GNNs) to effectively combine the strengths of multiple clustering models and enhance the clustering performance. The ensemble models are designed to handle diverse data characteristics, while the deep learning components capture complex non-linear relationships within the data. GNNs are employed to derive the final clustering outcomes by preserving spatial and temporal dependencies, making the approach well-suited for complex multidimensional spatiotemporal data. Experimental results from three real-world multivariate spatiotemporal data demonstrate the effectiveness of HEDGTC in accurately clustering and analyzing spatiotemporal patterns, outperforming state of the art ensemble models as well as traditional and individual deep clustering methods in terms of clustering performance and accuracy. The proposed method offers a robust framework for a wide range of applications, including climate modeling, geospatial analysis, and dynamic system forecasting. Francis Ndikum Nji, Omar Faruque, Mostafa Cham, Vandana Pursnani Janeja, Jianwu Wang 0001 |
IEEE Big Data | 4 |
| 2024 | CMAD: Advancing Understanding of Geospatial Clusters of Anomalous Melt Events in Sea Ice ExtentabstractTraditional statistical analyses do not reveal the spatial locations and the temporal occurrences of clusters of anomalous events that are responsible for a significant loss of sea ice extent. To address this problem, we present a novel method named Convolution Matrix Anomaly Detection (CMAD). The onset and progression of clusters of anomalous melting events over the Antarctic Sea ice are studied as loss in sea ice extent, which are essentially negative values, where the traditional convolutional operation of the Convolutional Neural Network (CNN) approach is ineffective. CMAD is based on an inverse max pooling concept in the convolutional operation of CNN to address this gap. CMAD is developed to offer a solution without using a neural network, and unlike a full CNN, it doesn't require any training or testing processes. Satellite images are utilized to establish the loss in the Antarctic region. Our analysis shows that anomalous melting patterns have significantly affected the Weddell and the Ross Sea regions more than any other regions of the Antarctic, consistent with the largest disappearance in sea ice extent over these two regions. These findings bolster the applicability of the inverse max pooling based CMAD in detecting the spatiotemporal evolution of clusters of anomalous melting events over the Antarctic region. The anomalous melting process was first noticed along the outer boundary of the sea ice extent in early September 2022 and gradually engulfed the entire sea ice region by February 2023 -in tandem with the scientific literature. These findings indicate that there is a necessity to delve deeper into the role of the anomalous melting process on sea ice retreat for a better understanding of the sea ice retreat process. The nature of the problem is to detect clusters of contiguous grids of anomalous melting events rather than detecting discrete grid points. CMAD's ability to perform both data clustering and anomaly detection via the pooling operations allows for a more comprehensive analysis of sea ice melt patterns, facilitating the pinpointing of areas with potentially significant melt events. This method has the potential to apply in other fields of study where anomalous events are detected in clusters. The inverse max pooling concept has successfully detected clusters of anomalous events in sea ice and demonstrated the capability to detect anomalies with 87% accuracy in benchmark data. In contrast to well-established conventional methods such as DBSCAN, HDBSCAN, K-Means, Bisecting K-Means, BIRCH, Agglomerative Clustering, OPTICS, and Gaussian Mixtures, when applied to dynamic multidimensional data, CMADBenchmark (which is a variation of CMAD) exhibits superior capabilities in detecting extreme events. The comparative analysis reveals that CMADBenchmark outperforms these traditional approaches, showcasing its heightened sensitivity and efficacy in capturing significant variations within evolving multidimensional datasets over time. This heightens the detection accuracy positions of CMAD as a valuable tool for discerning extreme events in the context of dynamic and changing multidimensional data. Maloy Kumar Devnath, Sudip Chakraborty, Vandana Pursnani Janeja |
SIGSPATIAL/GIS | 3 |
| 2024 | Interactive Assessment of Variances of High-Resolution Model Features in Digital Twin SimulationsabstractPrior to the deployment of expensive instruments into orbit, spatio-temporal digital twin systems modeling the whole earth are used to study the efficacy of these instruments. However, we need to make sure that the simulated instruments have realistic characteristics (to reflect the physics of the atmosphere and limits of the instrument itself) in order for the results of the digital twin to be robust and usable. If these simulations are done accurately, the instrument can be deployed, leading to more accurate weather forecasts and climate research. This demonstration system validates the simulations, specifically the realism of remotely sensed observations. The digital twin system is a low-cost way to improve instrument design used in meteorological and climatological research. The primary goal is to show how atmospheric data can improve the development and validation of new observational systems for meteorology and climate science. We have developed an interactive variability study system that uses a dynamic platform to visualize, assess, and grasp complex atmospheric dynamics. The dashboard is built using Python for backend operations and integrates tools such as the Streamlit framework for quick web application development and the Folium library for advanced geospatial visualizations. This dashboard acts as a bridge between advanced atmospheric modeling and spatio-temporal digital twin applications, showcasing the substantial benefits of integrating comprehensive model outputs into the simulation of observational systems. Chhaya Kulkarni, Nikki Privé, Vandana Pursnani Janeja |
SIGSPATIAL/GIS | 3 |
| 2024 | Discovery of multi-domain spatiotemporal associations
Prathamesh Walkikar, Bayu Adhi Tama, Vandana Pursnani Janeja |
GeoInformatica | 4 |
| 2017 | PhenomenaAssociater: Linking Multi-domain Spatio-Temporal Datasets
Prathamesh Walkikar, Vandana Pursnani Janeja |
DASFAA (2) | 2 |
| 2016 | Iterative unified clustering in big dataabstractWe propose a novel iterative unified clustering algorithm for data with both continuous and categorical variables, in the big data environment. Clustering is a well-studied problem and finds several applications. However, none of the big data clustering works discuss the challenge of mixed attribute datasets, with both categorical and continuous attributes. We study an application in the health care domain namely Case Based Reasoning (CBR), which refers to solving new problems based on solutions to similar past problems. This is particularly useful when there is a large set of clinical records with several types of attributes, from which similar patients need to be identified. We go one step further and include the genomic components of patient records to enhance the CBR discovery. Thus, our contributions in this paper spans across the big data algorithmic research and a key contribution to the domain of heath care information technology research. First, our clustering algorithm deals with both continuous and categorical variables in the data; second, our clustering algorithm is iterative where it finds the clusters which are not well formed and iteratively drills down to form well defined clusters at the end of the process; third we provide a novel approach to CBR across clinical and genomic data. Our research has implications for clinical trials and facilitating precision diagnostics in large and heterogeneous patient records. We present extensive experimental results to show the efficacy or our approach. Vasundhara Misal, Vandana Pursnani Janeja, Sai C. Pallaprolu, Yelena Yesha, Raghu Chintalapati |
IEEE BigData | 2 |
| 2016 | Label propagation in big data to detect remote access TrojansabstractRemote Access Trojans (RATs) provide cyber criminals with unlimited access to infected endpoints. Using the victim's access privileges, they can access and steal sensitive business and personal data including intellectual property and, personally identifiable information. However due to attack evolution, targeted attacks utilize modified versions of known signatures, which means that IDS rules that only match the known signature can be bypassed. In this paper, we propose a semi supervised approach that uses ensemble based label propagation to discover infected RAT packets in large unlabeled data. Our approach is trained on a small sample of labeled instances that usually characterize massive network datasets. Our approach is implemented using Apache Spark where we are able to demonstrate the effective discovery of such Trojans in massive amounts of data. We compare our approach to traditional signature based intrusion detection systems and clearly show that our approach is promising in the domain of cyber security in predicting large sets of unlabeled data using few labeled samples. Sai C. Pallaprolu, Josephine M. Namayanja, Vandana Pursnani Janeja, C. T. Sai Adithya |
IEEE BigData | 3 |
| 2015 | Unified framework for clinical data analytics (U-CDA)abstractIn spite of significant progress in the area of data management and integration, heterogeneous nature of clinical data makes it challenging to develop a unified view of clinical data. Therefore, a central question we are trying to address is how we can utilize data analytics to discover insightful knowledge from the scattered & large amount of clinical data to simplify clinical decision making. We propose a Unified Framework for Clinical Data Analytics (U-CDA) for mining large amounts of heterogeneous data to build enhanced clinical data analytics system. The proposed framework (U-CDA) in this paper integrates relevant clinical data from structured and unstructured data sources such as electronic health records, legacy health information system databases, clinical notes, public registries, and genomic datasets after applying necessary cleansing and transformations. It further uses intelligent, versatile data analytics engine to analyze clinical data. Jay Gholap, Vandana Pursnani Janeja, Yelena Yesha |
IEEE BigData | 2 |
| 2015 | SQL-like big data environments: Case study in clinical trial analyticsabstractBig Data deals with enormous volumes of complex and exponentially growing data sets from multiple sources. With rapid growth in technology, we are now able to generate immense amount of data in almost any field imaginable including physical, biological and biomedical sciences. With the diversity and amount of data in health care industry there is an increasing need to evaluate the components in big data frameworks and gauge their adaptability to analytics techniques. However, a key step in adapting big data tools is the portability of relational databases to big data environment. Since SQL is considered to be the de-facto language for interactive queries, in this paper, we evaluate the performance of SQL-like big data solutions for the portability of existing relational databases. Our work focuses on benchmarking multiple SQL-like big data technologies over Hadoop based distributed file system (HDFS) for Study Data Tabulation Model (SDTM) used in clinical trial databases for improving the efficiency of research in clinical trials. We use publically available clinical trial data (from National Institute on Drug Abuse (NIDA)), which follows SDTM, as a test bed to measure key parameters like usability, adaptability, modularity, robustness and efficiency of these solutions. With the intention to demonstrate how current clinical trial functionality can be replicated on a big data backend with high SQL-like functionality, we evaluate several types of ad-hoc SQL queries. Akshay Grover, Jay Gholap, Vandana Pursnani Janeja, Yelena Yesha, Raghu Chintalapati, Harsh Marwaha, Kunal Modi |
IEEE BigData | 3 |
| 2015 | STenSr: Spatio-temporal tensor streams for anomaly detection and pattern discovery
Aryya Gangopadhyay, Vandana Pursnani Janeja |
Knowl. Inf. Syst. | 3 |
| 2014 | B-dids: Mining anomalies in a Big-distributed Intrusion Detection SystemabstractThe focus of this paper is to present the architecture of a Big-distributed Intrusion Detection System (B-dIDS) to discover multi-pronged attacks which are anomalies existing across multiple subnets in a distributed network. The B-dIDS is composed of two key components: a big data processing engine and an analytics engine. The big data processing is done through HAMR, which is a next generation in-memory MapReduce engine. HAMR has reported high speedups over existing big data solutions across several analytics algorithms. The analytics engine comprises a novel ensemble algorithm, which extracts training data from clusters of the multiple IDS alarms. The clustering is utilized as a preprocessing step to re-label the datasets based on their high similarity to known potential attacks. The overall aim is to predict multi-pronged attacks that are spread across multiple subnets but can be missed if not evaluated in an integrated manner. Vandana Pursnani Janeja, Ali Azari, Josephine M. Namayanja, Brian Heilig |
IEEE BigData | 1 |
| 2014 | Change detection in temporally evolving computer networks: A big data frameworkabstractThe focus of this paper is to utilize a big data framework to characterize the behavior of central nodes over time in order to detect changes in large evolving computer networks. Changes in large evolving networks may indicate potential events such as cyber attacks, network failures or major shifts in network usage due to current events. Our approach entails the use of big data processing techniques such as MapReduce, to monitor central nodes to determine Consistency and Inconsistency (CoIn) in their availability and degree centrality across time periods. We also identify the Time Periods of Change (TPC) associated with CoIn. We present experimental results using real world internet traffic trace data, which indicates the potential of our approach to efficiently combine distributed processing techniques such as MapReduce to identify patterns of deviations. Josephine M. Namayanja, Vandana Pursnani Janeja |
IEEE BigData | 2 |
| 2014 | Mining trajectories of moving dynamic spatio-temporal regions in sensor datasets
Michael P. McGuire, Vandana Pursnani Janeja, Aryya Gangopadhyay |
Data Min. Knowl. Discov. | 2 |
| 2014 | Human perspective to anomaly detection for cybersecurity
Vandana Pursnani Janeja |
J. Intell. Inf. Syst. | 2 |
| 2013 | Multi-domain anomaly detection in spatial datasets
Vandana Pursnani Janeja, Revathi Palanisamy |
Knowl. Inf. Syst. | 1 |
| 2011 | Characterizing sensor datasets with multi-granular spatio-temporal intervalsabstractData from sensors and sensor networks are being collected at astronomical rates. This results in a massive dataset that is increasingly difficult to navigate to find interesting time periods where the spatial pattern of a process changes. The ability to navigate to such areas can lead to new knowledge about the factors that contribute to a spatio-temporal process. This paper proposes a method to automatically characterize sensor datasets based on a measure of spatial change over time resulting in a set of multi-granular spatio-temporal intervals. The resulting intervals can be used to focus knowledge discovery tasks at multiple temporal granularities within the dataset. Furthermore, the intervals enable a drill-down-style analysis where events of varying magnitudes can be identified within each granularity. Experiments were performed on a real-world dataset measuring NEXRAD precipitation accumulation. The results show that the multi-granular spatio-temporal intervals identify interesting time periods in the dataset as evidenced by naturally occurring events. Michael P. McGuire, Vandana Pursnani Janeja, Aryya Gangopadhyay |
GIS | 2 |
| 2011 | Anomalous Window Discovery for Linear Intersecting PathsabstractThe focus of this paper is to discover anomalous windows in linear intersecting paths. Anomalous windows are the contiguous groupings of data points. A linear path refers to a path represented by a line with a single dimensional spatial coordinate marking an observation point. In this paper, we propose an approach for discovering anomalous windows using a class of algorithms based on scan statistics, specifically 1) an Order invariant algorithm using Scan Statistics for Linear Intersecting Paths (SSLIP), 2) Brute force-SSLIP (BF-SSLIP), and 3) Central Brute Force-SSLIP (CBF-SSLIP). We further present two efficient variants of SSLIP: SSLIP* which employs a upper bound on the scan window size, and SSLIP-Acc, which adopts an accelerator function to speed up the scan process. The proposed approach for discovering anomalous windows along linear paths comprises the following distinct steps: 1) Cross Path Discovery: where we identify a subset of intersecting paths to be considered, 2) Anomalous Window Discovery: where we outline the various algorithms for the traversal of the cross paths to identify varying size directional windows along the paths. For identifying an anomalous window, an unusualness metric is computed, in the form of a likelihood ratio to indicate the degree of unusualness of this window with respect to the rest of the data. We identify the window with the highest likelihood ratio as our anomalous window, and 3) Monte Carlo Simulations: to ascertain whether this window is truly anomalous and not merely random occurrence, we perform hypothesis testing by computing a p-value using Monte Carlo Simulations. We present extensive experimental results in real world accident data sets for various highways with known issues (code and data available from [32], [27]). Additionally, we also perform comparisons with current approaches [18], [34] to show the efficacy of our approach. Our results show that our approach indeed is effective in identifying anomalous traffic accident windows along multiple intersecting highways. Vandana Pursnani Janeja |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2010 | Spatial neighborhood based anomaly detection in sensor datasets
Vandana Pursnani Janeja, Nabil R. Adam, Vijayalakshmi Atluri, Jaideep Vaidya |
Data Min. Knowl. Discov. | 1 |
| 2009 | Temporal Neighborhood Discovery Using Markov ModelsabstractTemporal data, which is a sequence of data tuples measured at successive time instances, is typically very large. Hence instead of mining the entire data, we are interested in dividing the huge data into several smaller intervals of interest which we call temporal neighborhoods. In this paper we propose an approach to generate temporal neighborhoods through unequal depth discretization. We describe two novel algorithms (a) similarity based merging (SMerg) and, (b) stationary distribution based merging (StMerg). These algorithms are based on the robust framework of Markov models and the Markov stationary distribution respectively. We identify temporal neighborhoods with distinct demarcations based on unequal depth discretization of the data. We discuss detailed experimental results in both synthetic and real world data. Specifically we show (i) the efficacy of our approach through precision and recall of labeled bins, (ii) the ground truth validation in real world datasets and, (iii) knowledge discovery in the temporal neighborhoods such as global anomalies. Our results indicate that we are able to identify valuable knowledge based on our ground truth validation from real world traffic data. Sandipan Dey, Vandana Pursnani Janeja, Aryya Gangopadhyay |
ICDM | 2 |
| 2009 | Anomalous window discovery through scan statistics for linear intersecting paths (SSLIP)abstractAnomalous windows are the contiguous groupings of data points. In this paper, we propose an approach for discovering anomalous windows using Scan Statistics for Linear Intersecting Paths (SSLIP). A linear path refers to a path represented by a line with a single dimensional spatial coordinate marking an observation point. Our approach for discovering anomalous windows along linear paths comprises of the following distinct steps: (a) Cross Path Discovery: where we identify a subset of intersecting paths to be considered, (b) Anomalous Window Discovery: where we outline three order invariant algorithms, namely SSLIP, Brute Force-SSLIP and Central Brute Force-SSLIP, for the traversal of the cross paths to identify varying size directional windows along the paths. For identifying an anomalous window we compute an unusualness metric, in the form of a likelihood ratio to indicate the degree of unusualness of this window with respect to the rest of the data. We identify the window with the highest likelihood ratio as our anomalous window, and (c) Monte Carlo Simulations: to ascertain whether this window is truly anomalous and not just a random occurrence we perform hypothesis testing by computing a p-value using Monte Carlo Simulations. We present extensive experimental results in real world accident datasets for various highways with known issues(code and data available from [27], [21]). Our results show that our approach indeed is effective in identifying anomalous traffic accident windows along multiple intersecting highways. Vandana Pursnani Janeja |
KDD | 2 |
| 2009 | Discretized Spatio-Temporal Scan WindowabstractThe focus of this paper is the discovery of anomalous spatio-temporal windows. We propose a Discretized Spatio-Temporal Scan Window approach to address the question of how we can treat Space and Time together without compromising on the properties of each and their impact on each other. In doing so we discover anomalous Spatio-Temporal windows, identify at what point in time the window changes, identify the spatial patterns of change over time and identify a spatial extent in time which is completely deviant with respect to the rest of the anomalous spatio-temporal windows. None of the current approaches address all these issues in combination. Subsequently we perform experiments on several real world datasets to validate our approach while comparing with the established approach of discovering a cylindrical spatio-temporal Scan window. Seyed H. Mohammadi, Vandana Pursnani Janeja, Aryya Gangopadhyay |
SDM | 2 |
| 2008 | Random Walks to Identify Anomalous Free-Form Spatial Scan WindowsabstractOften, it is required to identify anomalous windows over a spatial region that reflect unusual rate of occurrence of a specific event of interest. A spatial scan statistic-based approach essentially considers a scan window and computes the statistic of a parameter(s) of interest, and identifies anomalous windows by moving the scan window in the region. While this approach has been successfully employed in identifying anomalous windows, earlier proposals adopting spatial scan statistic suffer from two limitations: (1) In general, the scan window should be of a regular shape (e.g., circle, rectangle, cylinder). Thus, most approaches are capable of identifying anomalous windows of fixed shapes only. However, the region of anomaly, in general, is not necessarily of a regular shape. Recent proposals to identify windows of irregular shapes identify windows much larger than the true anomalies or penalize large-sized windows. (2) These techniques take into account autocorrelation among spatial data but not spatial heterogeneity. As a result, they often result in inaccurate anomalous windows. To address these limitations, in this paper, we propose a random-walk-based Free-Form Spatial Scan Statistic (FS3). We construct a weighted Delaunay nearest neighbor (WDNN) graph to capture both spatial autocorrelation and heterogeneity. We then use random walks to identify natural free-form scan windows that are not restricted to a predefined shape. We use spatial scan statistics to identify anomalous windows and prove that they are not random but indeed are formed as a result of an anomaly. Application of FS3on real data sets has shown that it can identify more refined anomalous windows with better likelihood ratio of it being an anomaly than those identified by earlier spatial scan statistic approaches. Vandana Pursnani Janeja, Vijayalakshmi Atluri |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2005 | FS3: A Random Walk Based Free-Form Spatial Scan Statistic for Anomalous Window DetectionabstractOften, it is required to identify anomalous windows over a spatial region that reflect unusual rate of occurrence of a specific event of interest. A spatial scan statistic essentially considers a scan window, and identifies anomalous windows by moving the scan window in the region. While spatial scan statistic has been successful, earlier proposals suffer from two limitations: (i) They restrict the scan window to be of a regular shape (e.g., circle, rectangle, cylinder). However, the region of anomaly, in general, is not necessarily of a regular shape. (ii) They take into account autocorrelation among spatial data, but not spatial heterogeneity. As a result, they often result in inaccurate anomalous windows. To address these limitations, we propose a random walk based free-form spatial scan statistic (FS/sup 3/). Application of FS/sup 3/ on real datasets has shown that it can identify more refined anomalous windows with better likelihood ratio of it being an anomaly, than those identified by earlier spatial scan statistic approaches. Vandana Pursnani Janeja, Vijayalakshmi Atluri |
ICDM | 1 |