James M. Kang

dblp:23/6499 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 12 · 5 first-authorArtificial intelligence and machine learning · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
8 papers
Data mining · 52% Spatial and temporal data management · 41% Data stream processing · 4%
Computer networks
1 paper
Internet of things and sensor networks · 100%

Topics — the 20 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining › structured data mining
spatial data mining
0.212016
Ring-Shaped Hotspot Detection · IEEE Trans. Knowl. Data Eng. 2016
Data mining › structured data mining › spatial data mining
spatial scan statistic
0.212016
Ring-Shaped Hotspot Detection · IEEE Trans. Knowl. Data Eng. 2016
Data mining
clustering
0.212014
A K-Main Routes Approach to Spatial Network Activity Summarization · IEEE Trans. Knowl. Data Eng. 2014
Data mining › clustering
k-means clustering
0.212014
A K-Main Routes Approach to Spatial Network Activity Summarization · IEEE Trans. Knowl. Data Eng. 2014
Spatial and temporal data management
spatial network
0.212014
A K-Main Routes Approach to Spatial Network Activity Summarization · IEEE Trans. Knowl. Data Eng. 2014
Spatial and temporal data management
spatio-temporal query processing
0.212014
Lagrangian Approaches to Storage of Spatio-Temporal Network Datasets · IEEE Trans. Knowl. Data Eng. 2014
Spatial and temporal data management › spatial query processing › nearest neighbor query
reverse nearest neighbor query
0.222010
Incremental and General Evaluation of Reverse Nearest Neighbors · IEEE Trans. Knowl. Data Eng. 2010
Continuous Evaluation of Monochromatic and Bichromatic Reverse Nearest Neighbors · ICDE 2007
Spatial and temporal data management
spatial query processing
0.222010
Incremental and General Evaluation of Reverse Nearest Neighbors · IEEE Trans. Knowl. Data Eng. 2010
Continuous Evaluation of Monochromatic and Bichromatic Reverse Nearest Neighbors · ICDE 2007
Data stream processing
continuous query processing
0.112010
Incremental and General Evaluation of Reverse Nearest Neighbors · IEEE Trans. Knowl. Data Eng. 2010
Data mining
anomaly detection
0.112008
Discovering Flow Anomalies: A SWEET Approach · ICDM 2008
Internet of things and sensor networks › environmental sensing
environmental sensor network
0.112008
Discovering Flow Anomalies: A SWEET Approach · ICDM 2008
Data mining › clustering
spatial clustering
0.112016
Ring-Shaped Hotspot Detection · IEEE Trans. Knowl. Data Eng. 2016
Information retrieval › evaluation › online evaluation
continuous evaluation
0.112007
Continuous Evaluation of Monochromatic and Bichromatic Reverse Nearest Neighbors · ICDE 2007
Data mining › structured data mining › spatial data mining
spatial co-location pattern mining
0.112007
Zonal Co-location Pattern Discovery with Dynamic Parameters · ICDM 2007
Spatial and temporal data management
spatial indexing
0.112007
Zonal Co-location Pattern Discovery with Dynamic Parameters · ICDM 2007
Spatial and temporal data management
shortest path computation
0.112014
A K-Main Routes Approach to Spatial Network Activity Summarization · IEEE Trans. Knowl. Data Eng. 2014
Storage systems › i/o optimization
disk i/o optimization
0.112014
Lagrangian Approaches to Storage of Spatio-Temporal Network Datasets · IEEE Trans. Knowl. Data Eng. 2014
Logic in computer science
query evaluation
0.012010
Incremental and General Evaluation of Reverse Nearest Neighbors · IEEE Trans. Knowl. Data Eng. 2010
Data mining
pattern mining
0.012007
Zonal Co-location Pattern Discovery with Dynamic Parameters · ICDM 2007
Computational geometry
voronoi diagram
0.012007
Continuous Evaluation of Monochromatic and Bichromatic Reverse Nearest Neighbors · ICDE 2007

Methods — techniques the papers use, named apart from their topics

dual grid based pruning · 0.4space-filling curves · 0.4lagrangian family set · 0.4pruning · 0.4best enclosing ring refining · 0.2filter-and-refine · 0.2computational complexity analysis · 0.2statistical significance test · 0.2network voronoi · 0.2divide-and-conquer · 0.2smart window enumeration · 0.1voronoi diagram · 0.1incremental monitoring · 0.1
YearPublicationVenuePosition
2016 Ring-Shaped Hotspot Detection
abstract
Given a set of activity points (e.g., crime, disease locations), Ring-Shaped Hotspot Detection (RHD) finds ring-shaped areas where the concentration of activities inside is significantly higher than that outside. RHD is societally important for applications such as environmental criminology, epidemiology, and biology to investigate evasive patterns. RHD is computationally challenging because of the large number of candidate rings, non-monotonic interest measure, and cost of the statistical significance test. Previous approaches (e.g., spatial scan statistics tools) focus on simply-connected shaped areas (e.g., circles, rectangles) and can not detect statistically significant rings. In this paper, a novel algorithm, DGPLMR, is proposed to discover statistically significant ring-shaped hotspots based on the ideas of dual grid based pruning and best enclosing ring refining. Theoretical evaluation proves that the proposed approach is a correct approach (i.e., all outputs satisfy input thresholds) to detect ring-shaped hotspots. Case study on real disease data shows that the proposed approach finds ring-shaped hotspots which were not detected by the existing techniques. Cost analysis and experimental results on synthetic data show that the proposed approach with algorithmic refinements yields substantial computational savings.
Emre Eftelioglu, Shashi Shekhar 0001, James M. Kang, Christopher Farah
IEEE Trans. Knowl. Data Eng.3
2014 Ring-Shaped Hotspot Detection: A Summary of Results
abstract
Given a collection of geo-located activities (e.g., Crime reports), ring-shaped hotspot detection (RHD) finds rings, where concentration of activities inside the ring is much higher than outside. RHD is important for the applications such as crime analysis, where it may focus the search for crime source's location, e.g. The home of a serial criminal. RHD is challenging because of the large number of candidate rings and the high computational cost of the statistical significance test. Previous statistically significant hotspot detection techniques (e.g., Sat Scan) identify circular/rectangular areas, but can not discover rings. This paper proposes a dual grid based pruning (DGP) approach to detect ring-shaped hotspots. A case study on real crime data confirms that DGP detects novel ring-shaped regions, regions that go undetected by Sat Scan. Experiments show that DGP improves the computational cost of a naive approach substantially.
Emre Eftelioglu, Shashi Shekhar 0001, Dev Oliver, Xun Zhou 0001, Michael R. Evans, Yiqun Xie, James M. Kang, Renee Laubscher, Christopher Farah
ICDM7
2014 A K-Main Routes Approach to Spatial Network Activity Summarization
abstract
Data summarization is an important concept in data mining for finding a compact representation of a dataset. In spatial network activity summarization (SNAS), we are given a spatial network and a collection of activities (e.g., pedestrian fatality reports, crime reports) and the goal is to find k shortest paths that summarize the activities. SNAS is important for applications where observations occur along linear paths such as roadways, train tracks, etc. SNAS is computationally challenging because of the large number of k subsets of shortest paths in a spatial network. Previous work has focused on either geometry or subgraph-based approaches (e.g., only one path), and cannot summarize activities using multiple paths. This paper proposes a K-Main Routes (KMR) approach that discovers k shortest paths to summarize activities. KMR generalizes K-means for network space but uses shortest paths instead of ellipses to summarize activities. To improve performance, KMR uses network Voronoi, divide and conquer, and pruning strategies. We present a case study comparing KMR's network-based output (i.e., shortest paths) to geometry-based outputs (e.g., ellipses) on pedestrian fatality data. Experimental results on synthetic and real data show that KMR with our performance-tuning decisions yields substantial computational savings without reducing summary path coverage.
Dev Oliver, Shashi Shekhar 0001, James M. Kang, Renee Laubscher, Veronica Carlan, Abdussalam Bannur
IEEE Trans. Knowl. Data Eng.3
2014 Lagrangian Approaches to Storage of Spatio-Temporal Network Datasets
abstract
Given a spatio-temporal network (STN) and a set of STN operations, the goal of the Storing Spatio-Temporal Networks (SSTN) problem is to produce an efficient method of storing STN data that minimizes disk I/O costs for given STN operations. The SSTN problem is important for many societal applications, such as surface and air transportation management systems. The problem is NP hard, and is challenging due to an inherently large data volume and novel semantics (e.g., Lagrangian reference frame). Related works rely on orthogonal partitioning approaches (e.g., snapshot and longitudinal) and incur excessive I/O costs when performing common STN queries. Our preliminary work proposed a non-orthogonal partitioning approach in which we optimized the LGetOneSuccessor() operation that retrieves a single successor for a given node on STN. In this paper, we provide a method to optimize the LGetAllSuccessors() operation, which retrieves all successors for a given node on a STN. This new approach uses the concept of a Lagrangian Family Set (LFS) to model data access patterns for STN queries. Experimental results using real-world road and flight traffic datasets demonstrate that the proposed approach outperforms prior work for LGetAllSuccessors() computation workloads.
KwangSoo Yang, Michael R. Evans, Venkata M. V. Gunturi, James M. Kang, Shashi Shekhar 0001
IEEE Trans. Knowl. Data Eng.4
2011 Tipping Points, Butterflies, and Black Swans: A Vision for Spatio-temporal Data Mining Analysis
James M. Kang, Daniel L. Edwards
SSTD1
2010 A Lagrangian approach for storage of spatio-temporal network datasets: a summary of results
abstract
Given a set of operators and a spatio-temporal network, the goal of the Storing Spatio-Temporal Networks (SSTN) problem is to produce an efficient data storage method that minimizes disk I/O access costs. Storing and accessing spatio-temporal networks is increasingly important in many societal applications such as transportation management and emergency planning. This problem is challenging due to strains on traditional adjacency list representations when storing temporal attribute values from the sizable increase in length of the time-series. Current approaches for the SSTN problem focus on orthogonal partitioning (e.g., snapshot, longitudinal, etc.), which may produce excessive I/O costs when performing traversal-based spatio-temporal network queries (e.g., route evaluation, arrival time prediction, etc) due to the desired nodes not being allocated to a common page. We propose a Lagrangian-Connectivity Partitioning (LCP) technique to efficiently store and access spatio-temporal networks that utilizes the interaction between nodes and edges in a network. Experimental evaluation using the Minneapolis, MN road network showed that LCP outperforms traditional orthogonal approaches.
Michael R. Evans, KwangSoo Yang, James M. Kang, Shashi Shekhar 0001
GIS3
2010 Incremental and General Evaluation of Reverse Nearest Neighbors
abstract
This paper presents a novel algorithm for Incremental and General Evaluation of continuous Reverse Nearest neighbor queries (IGERN, for short). The IGERN algorithm is general in that it is applicable for both continuous monochromatic and bichromatic reverse nearest neighbor queries. This problem is faced in a number of applications such as enhanced 911 services and in army strategic planning. A main challenge in these problems is to maintain the most up-to-date query answers as the data set frequently changes over time. Previous algorithms for monochromatic continuous reverse nearest neighbor queries rely mainly on monitoring at the worst case of six pie regions, whereas IGERN takes a radical approach by monitoring only a single region around the query object. The IGERN algorithm clearly outperforms the state-of-the-art algorithms in monochromatic queries. We also propose a new optimization for the monochromatic IGERN to reduce the number of nearest neighbor searches. Furthermore, a filter and refine approach for IGERN (FR-IGERN) is proposed for the continuous evaluation of bichromatic reverse nearest neighbor queries which is an optimized version of our previous approach. The computational complexity of IGERN and FR-IGERN is presented in comparison to the state-of-the-art algorithms in the monochromatic and bichromatic cases. In addition, the correctness of IGERN and FR-IGERN in both the monochromatic and bichromatic cases, respectively, are proved. Extensive experimental analysis using synthetic and real data sets shows that IGERN and FR-IGERN is efficient, is scalable, and outperforms previous techniques for continuous reverse nearest neighbor queries.
James M. Kang, Mohamed F. Mokbel, Shashi Shekhar 0001, Tian Xia 0001
IEEE Trans. Knowl. Data Eng.1
2009 Discovering Teleconnected Flow Anomalies: A Relationship Analysis of Dynamic Neighborhoods (RAD) Approach
James M. Kang, Shashi Shekhar 0001, Michael Henjum, Paige J. Novak, William A. Arnold
SSTD1
2009 Spatio-Temporal Sensor Graphs (STSG): A data model for the discovery of spatio-temporal patterns
abstract
Developing a model that facilitates the representation and knowledge discovery on sensor data presents many challenges. With sensors reporting data at a very high frequency, resulting in large volumes of data, there is a need for a model that is memory efficient. Since sensor data is spatio-tempora l in nature, the model must also support the time dependence of the data. Balancing the conflicting requirements of simplicity, expressiveness and storage efficiency is challenging. The model should also provide adequate support for the formulation of efficient algorithms for knowledge discovery. Though spatio-temporal data can be modeled using time expanded graphs, this model replicates the entire graph across time instants, resulting in high storage overhead and computationally expensive algorithms. In this paper, we propose Spatio-Temporal Sensor Graphs (STSG) to model sensor data at the conceptual. logical and physical levels. This model allows the properties of edges and nodes to be modeled as a time series of measurement data. Data at each instant would consist of the measured value and the expected error. Also, we evaluate the model using methods to find interesting patterns such as growing hotspots in sensor data and present analytical comparison of the algorithms with methods based on existing models.
Betsy George, James M. Kang, Shashi Shekhar 0001
Intell. Data Anal.2
2009 Context inclusive function evaluation: a case study with EM-based multi-scale multi-granular image classification
Vijay Gandhi, James M. Kang, Shashi Shekhar 0001, Junchang Ju, Eric D. Kolaczyk, Sucharita Gopal
Knowl. Inf. Syst.2
2008 Discovering Flow Anomalies: A SWEET Approach
abstract
Given a percentage-threshold and readings from a pair of consecutive upstream and downstream sensors, flow anomaly discovery identifies dominant time intervals where the fraction of time instants of significantly mis-matched sensor readings exceed the given percentage-threshold. Discovering flow anomalies (FA) is an important problem in environmental flow monitoring networks and early warning detection systems for water quality problems. However, mining FAs is computationally expensive because of the large (potentially infinite) number of time instants of measurement and potentially long delays due to stagnant (e.g. lakes) or slow moving (e.g. wetland) water bodies between consecutive sensors. Traditional outlier detection methods (e.g. t-test) are suited for detecting transient FAs (i.e., time instants of significant mis-matches across consecutive sensors) and cannot detect persistent FAs (i.e., long variable time-windows with a high fraction of time instant transient FAs) due to a lack of a pre-defined window size. In contrast, we propose a Smart Window Enumeration and Evaluation of persistence-Thresholds (SWEET) method to efficiently explore the search space of all possible window lengths. Computation overhead is brought down significantly by restricting the start and end points of a window to coincide with transient FAs, using a smart counter and efficient pruning techniques. Experimental evaluation using a real dataset shows our proposed approach outperforms Nainodotve alternatives.
James M. Kang, Shashi Shekhar 0001, Christine Wennen, Paige J. Novak
ICDM1
2007 Continuous Evaluation of Monochromatic and Bichromatic Reverse Nearest Neighbors
abstract
This paper presents a novel algorithm for Incremental and General Evaluation of continuous Reverse Nearest neighbor queries (IGERN, for short). The IGERN algorithm is general as it is applicable for both the monochromatic and bichromatic reverse nearest neighbor queries. The incremental aspect of IGERN is achieved through determining only a small set of objects to be monitored. While previous algorithms for monochromatic queries rely mainly on monitoring six pie regions, IGERN takes a radical approach by monitoring only a single region around the query object. The IGERN algorithm clearly outperforms the state-of-the-art algorithms in monochromatic queries. In addition, the IGERN algorithm presents the first attempt for continuous evaluation of bichromatic reverse nearest neighbor queries. The computational complexity of IGERN is presented in comparison to the state-of-the-art algorithms in the monochromatic case and to the use of Voronoi diagrams for the bichromatic case. In addition, the correctness of IGERN in both the monochromatic and bichromatic cases are proved. Extensive experimental analysis shows that IGERN is efficient, is scalable, and outperforms previous techniques for continuous reverse nearest neighbor queries.
James M. Kang, Mohamed F. Mokbel, Shashi Shekhar 0001, Tian Xia 0001
ICDE1
2007 Zonal Co-location Pattern Discovery with Dynamic Parameters
abstract
Zonal co-location patterns represent subsets of featuretypes that are frequently located in a subset of space (i.e., zone). Discovering zonal spatial co-location patterns is an important problem with many applications in areas such as ecology, public health, and homeland defense. However, discovering these patterns with dynamic parameters (i.e., repeated specification of zone and interest measure values according to user preferences) is computationally complex due to the repetitive mining process. Also, the set of candidate patterns is exponential in the number of feature types, and spatial datasets are huge. Previous studies have focused on discovering global spatial co-location patterns with a fixed interest measure threshold. In this paper, we propose an indexing structure for co-location patterns and propose algorithms (Zoloc-Miner) to discover zonal colocation patterns efficiently for dynamic parameters. Extensive experimental evaluation shows our proposed approaches are scalable, efficient, and outperform na¨ive alternatives.
Mete Celik, James M. Kang, Shashi Shekhar 0001
ICDM2