Ashok Samal

dblp:52/4454 · DBLP profile ↗
← Back
18ranked-venue papers in the field
2as first author
3since 2021 · last 2024
0000-0002-4559-9454ORCID · corroborated

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 7Database Systems & Data Management · 5 (1 first)Data Mining & Knowledge Discovery · 4Big Data, Cloud & Distributed Data Systems · 1Knowledge Engineering, Semantic Web & Information Systems · 1 (1 first)
YearPublicationVenuePosition
2024 Discovering Localized Drivers of Unrest Events using Clustering and XGBoost
abstract
Social unrest, a multifaceted phenomenon that is influenced by a variety of interconnected factors, presents substantial obstacles to societal stability and governance. The comprehension of local nuances is frequently restricted by the analysis of drivers of unrest at broad geographic scales or the isolation of specific causes in traditional studies. This paper introduces the SCEIGE framework, which classifies unrest drivers into six essential categories: Socio-demographic, Cultural, Environmental, Infrastructural, Geographic, and Economic. This framework is designed to address these challenges. SCEIGE offers a comprehensive perspective on the fundamental causes of social unrest by modeling geographic spaces at fine resolutions and incorporating a wide range of variables. We further enhance this framework by introducing a novel clustering and machine learning methodology, SC-XG (SCEIGE Clustering with XGBoost), which organizes geographic regions according to SCEIGE patterns. SC-XG not only reveals the local drivers of unrest but also facilitates the predictive analysis of social unrest events. This paper also illustrates the effectiveness of high-resolution SCEIGE geo-rasters in analyzing social unrest and confirms that the drivers of unrest differ across regions, underscoring the necessity of a local-level understanding. We identify critical, region-specific unrest drivers and address the broader implications for predicting and mitigating social unrest globally by applying SC-XG to unrest patterns in India.
Dalton J. Hazelwood, Deepti Joshi, Ashok Samal, Leen-Kiat Soh
IEEE Big Data3
2023 A spatially-aware algorithm for location extraction from structured documents
Praval Sharma, Ashok Samal, Leen-Kiat Soh, Deepti Joshi
GeoInformatica2
2021 An information fusion approach for conflating labeled point-based time-series data
Zion Schell, Ashok Samal, Leen-Kiat Soh
GeoInformatica2
2017 SURGE: Social Unrest Reconnaissance GazEteer
abstract
Social Unrest Reconnaissance Gazetteer (or SURGE) is a Web-based application that provides an open system to visualize and integrate spatio-temporal data about social unrest events with related data layers in South Asia to facilitate data-driven as well as model-based investigations and analyses. Currently, the system displays eight categories of unrest, based primarily on the Global Database of Events, Language and Tone (GDELT) and the Global Terrorism Database (GTD). Users have the ability to select a single day or a range of dates along with the category of unrest they are interested to investigate. The users also have the option to normalize the raw event counts by population density. Additionally, the users can view infrastructure layers that facilitate or hinder the diffusion of unrest events (e.g., collated from an open GIS data-source: OpenStreetMap (www.openstreetmap.org)) and choropleth layers to display various socio-economic indicators (e.g., derived from global surveys and government census data such as the 2011 India census data (cenusindia.gov.in) and IPUMS Terra (data.terrapop.org)). Currently, SURGE displays unrest events for India, Pakistan and Bangladesh as heat map layers in multiple spatial resolutions. Challenges have involved geo-synchronization, data conversions, and displaying multiple layers of dense geospatial datasets. Future capabilities include automatic ingestion of raw data and standardizing levels of unrest using significant predictors.
Deepti Joshi, Sudeep Basnet, Hariharan Arunachalam, Leen-Kiat Soh, Ashok Samal, Shawn Ratcliff, Regina Werum
SIGSPATIAL/GIS5
2014 Using spatial data support for reducing uncertainty in geospatial applications
T. Hong, K. Hart, Leen-Kiat Soh, Ashok Samal
GeoInformatica4
2014 A dissimilarity function for geospatial polygons
Deepti Joshi, Leen-Kiat Soh, Ashok Samal
Knowl. Inf. Syst.3
2013 Spatio-temporal polygonal clustering with space and time as first-class citizens
Deepti Joshi, Ashok Samal, Leen-Kiat Soh
GeoInformatica2
2012 Redistricting Using Constrained Polygonal Clustering
abstract
Redistricting is the process of dividing a geographic area consisting of spatial units-often represented as spatial polygons-into smaller districts that satisfy some properties. It can therefore be formulated as a set partitioning problem where the objective is to cluster the set of spatial polygons into groups such that a value function is maximized [1]. Widely used algorithms developed for point-based data sets are not readily applicable because polygons introduce the concepts of spatial contiguity and other topological properties that cannot be captured by representing polygons as points. Furthermore, when clustering polygons, constraints such as spatial contiguity and unit distributedness should be strategically addressed. Toward this, we have developed the Constrained Polygonal Spatial Clustering (CPSC) algorithm based on the A* search algorithm that integrates cluster-level and instance-level constraints as heuristic functions. Using these heuristics, CPSC identifies the initial seeds, determines the best cluster to grow, and selects the best polygon to be added to the best cluster. We have devised two extensions of CPSC-CPSC* and CPSC*-PS-for problems where constraints can be soft or relaxed. Finally, we compare our algorithm with graph partitioning, simulated annealing, and genetic algorithm-based approaches in two applications-congressional redistricting and school districting.
Deepti Joshi, Leen-Kiat Soh, Ashok Samal
IEEE Trans. Knowl. Data Eng.3
2009 Density-based clustering of polygons
abstract
Clustering is an important task in spatial data mining and spatial analysis. We propose a clustering algorithm P-DBSCAN to cluster polygons in space. P-DBSCAN is based on the well established density-based clustering algorithm DBSCAN. In order to cluster polygons, we incorporate their topological and spatial properties in the process of clustering by using a distance function customized for the polygon space. The objective of our clustering algorithm is to produce spatially compact clusters. We measure the compactness of the clusters produced using P-DBSCAN and compare it with the clusters formed using DBSCAN, using the Schwartzberg index. We measure the effectiveness and robustness of our algorithm using a synthetic dataset and two real datasets. Results show that the clusters produced using P-DBSCAN have a lower compactness index (hence more compact) than DBSCAN.
Deepti Joshi, Ashok Samal, Leen-Kiat Soh
CIDM2
2009 A dissimilarity function for clustering geospatial polygons
abstract
The traditional point-based clustering algorithms when applied to geospatial polygons may produce clusters that are spatially disjoint due to their inability to consider various types of spatial relationships between polygons. In this paper, we propose to represent geospatial polygons as sets of spatial and non-spatial attributes. By representing a polygon as a set of spatial and non-spatial attributes we are able to take into account all the properties of a polygon (such as structural, topological and directional) that were ignored while using point-based representation of polygons, and that aid in the formation of high quality clusters. Based on this framework we propose a dissimilarity function that can be plugged into common state-of-the-art spatial clustering algorithms. The result is clusters of polygons that are more compact in terms of cluster validity and spatial contiguity. We show the effectiveness and robustness of our approach by applying our dissimilarity function on the traditional k-means clustering algorithm and testing it on a watershed dataset.
Deepti Joshi, Ashok Samal, Leen-Kiat Soh
GIS2
2009 Redistricting Using Heuristic-Based Polygonal Clustering
abstract
Redistricting is the process of dividing a geographic area into districts or zones. This process has been considered in the past as a problem that is computationally too complex for an automated system to be developed that can produce unbiased plans. In this paper we present a novel method for redistricting a geographic area using a heuristic-based approach for polygonal spatial clustering. While clustering geospatial polygons several complex issues need to be addressed - such as: removing order dependency, clustering all polygons assuming no outliers, and strategically utilizing domain knowledge to guide the clustering process. In order to address these special needs, we have developed the constrained polygonal spatial clustering (CPSC) algorithm that holistically integrates do-main knowledge in the form of cluster-level and instance-level constraints and uses heuristic functions to grow clusters. In order to illustrate the usefulness of our algorithm we have applied it to the problem of formation of unbiased congressional districts. Furthermore, we compare and contrast our algorithm with two other approaches proposed in the literature for redistricting, namely-graph partitioning and simulated annealing.
Deepti Joshi, Leen-Kiat Soh, Ashok Samal
ICDM3
2008 Computing information gain for spatial data support
abstract
Widespread use of GPS devices and explosion of remotely sensed geospatial images along with cheap storage devices has resulted in vast amounts of data. More recently, with the advent of wireless technology, a large number of sensor networks have been deployed to monitor many human, biological and natural processes. This poses a challenge in many data rich application domains. The problem now is how best to choose the datasets to solve specific problems. Some of the datasets may be redundant and their inclusion in analysis may not only be time consuming, but may lead to erroneous conclusions. We propose the concept of data support as the basis for efficient, cost-effective and intelligent use of geospatial data in order to reduce uncertainty in the analysis and consequently in the results. Data support is defined as the process of determining the information utility of a data source to help decide which one to include or exclude to improve cost-effectiveness in existing data analysis. In this article we use mutual information as the basis of computing data support. The concept of mutual information is defined in information theory as a measure to compute information gain or loss between two disjoint datasets. We use this to compute the optimal datasets in specific applications. The effectiveness of the approach is demonstrated using an application in the hydrological analysis domain.
Ashok Samal, Leen-Kiat Soh
GIS2
2008 Techniques for Computing Fitness of Use (FoU) for Time Series Datasets with Applications in the Geospatial Domain
Leen-Kiat Soh, Ashok Samal
GeoInformatica3
2006 Texture as the basis for individual tree identification
Ashok Samal, James R. Brandle
Inf. Sci.1
2005 Face Recognition Using Landmark-Based Bidimensional Regression
abstract
This paper studies how biologically meaningful landmarks extracted from face images can be exploited for face recognition using the bidimensional regression. Incorporating the correlation statistics of landmarks, this paper also proposes a new approach called eigenvalue weighted bidimensional regression. Complex principal component analysis is used for computing eigenvalues and removing correlation among landmarks. We evaluate our approach using two standard face databases: the Purdue AR and the NIST FERET. Experimental results show that the bidimensional regression is an efficient method to exploit geometry information of face images.
Jiazheng Shi, Ashok Samal, David Marx
ICDM2
2004 A feature-based approach to conflation of geospatial sources
abstract
A Geographic Information System (GIS) populated with disparate data sources has multiple and different representations of the same real-world object. Often, the type of information in these sources is different, and combining them to generate one composite representation has many benefits. The first step in this conflation process is to identify the features in different sources that represent the same real-world entity. The matching process is not simple, since the identified features from different sources do not always match in their location, extent, and description. We present a new approach to matching GIS features from disparate sources. A graph theoretic approach is used to model the geographic context and to determine the matching features from multiple sources. Experiments on implementation of this approach demonstrate its viability.
Ashok Samal, Sharad C. Seth, Kevin Cueto
Int. J. Geogr. Inf. Sci.1
1999 Cooperative Text and Line-Art Extraction from a Topographic Map
abstract
The black layer is digitized from a USGS topographic map digitized at 1000 dpi. The connected components of this layer are analyzed and separated into line art, text, and icons in two passes. The paired street casings are converted to polylines by vectorization and associated with street labels from the character recognition phase. The accuracy of character recognition is shown to improve by taking account of the frequently occurring overlap of line art with street labels. The experiments show that complete vectorization of the black line-layer bitmap is the major remaining problem.
George Nagy, Ashok Samal, Sharad C. Seth
ICDAR3
1995 A system for recognizing a large class of engineering drawings
abstract
We present a complete system for recognizing a large class of symbolic engineering drawings that includes flowcharts, chemical plant diagrams, and logic & electrical circuits. The output of the system, a netlist identifying the symbol types and interconnections, may be used for design verification or as a compact portable representation of the drawing. The automatic recognition task is done in two stages: (1) domain-independent rules segment symbols from connection lines in the preprocessed drawing image and (2) an understanding subsystem makes use of a set of domain-specific matchers to classify symbols and correct errors automatically. A graphical user interface is provided to correct residual errors interactively. The system has been tested on a large database of printed images drawn from four different domains.
Yuhong Yu, Ashok Samal, Sharad C. Seth
ICDAR2