Joonho Kwon

dblp:70/397 · DBLP profile ↗
← Back
11ranked-venue papers in the field
3as first author
3since 2021 · last 2022
0000-0002-8207-9415ORCID · corroborated

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 5Database Systems & Data Management · 3 (3 first)Data Mining & Knowledge Discovery · 1Information Retrieval & Web Search · 1Knowledge Engineering, Semantic Web & Information Systems · 1
YearPublicationVenuePosition
2022 Fairness Improvement Technology and Visualization Services for binary classification datasets collected on a batch basis
abstract
Since securing data is essential in the field of artificial intelligence, the importance of data collection and purification is steadily increasing. On the other hand, there are typical problems arising in the process of collecting data, such as the black bias problem in COMPAS and the fairness problem that exists within the data, such as the facial recognition bias problem. Accordingly, by designing and implementing a system that corrects fairness by removing bias existing in the dataset itself, we tried to help research on artificial intelligence models. The 'Fairness Improvement Technology and Visualization Services for binary classification datasets collected on a batch basis' consists of a subset generator that separates the initial binary classification datasets with unique values, a bias remover that removes bias by comparing and verifying each subset, and a visualization module that visualizes the corrected data as a web service. In addition, the proposed system in this paper presents a data-level research method that eliminates bias in the data itself without modifying additional learning or algorithms, and confirms the potential of the proposed system using COMPAS Data[1], Adult Census Income Data[2] as validation data.
Kyeongsu Byun, Joonho Kwon, Goo Kim
IEEE Big Data2
2022 X-FIST: Extended flood index for efficient similarity search in massive trajectory dataset
abstract
Similarity search tasks in big trajectory datasets often require tree-based indices to shorten the query time through early pruning of dissimilar trajectories early. However, tree-based indices have been outperformed by the learned index in skewed-distribution datasets of multidimensional point experimentally. The learned index performed faster because of its data distribution awareness and machine learning model-based prediction. Directly applying learned index to trajectories can lead to inefficient query performance due to repeating range queries according to the query trajectory length. Thus, we develop X-FIST, an extended Flood index to learn the Minimum Bounding Region of the trajectories and their sub-trajectories. In similarity search, X-FIST prunes dissimilar trajectories effectively independent to the query trajectory length. If the trajectory similarity distance function changes, X-FIST does not need to train new models of its Flood index. The experimental results on three real-world trajectory datasets demonstrate that our approach shortened query time in every distance function and produced better storage size reduction than the tree-based index and direct approach of learned index.
Hani Ramadhan, Joonho Kwon
Inf. Sci.2
2021 Enhancing Learned Index for A Higher Recall Trajectory K-Nearest Neighbor Search
abstract
Learned indices can significantly shorten the query response time of k-Nearest Neighbor search of points data. However, extending the learned index for k-Nearest Neighbor search of trajectory data may return incorrect results (low recall) and require longer pruning time. Thus, we introduce an enhancement for trajectory learned index which is a pruning step for a learned index to retrieve the k-Nearest Neighbors correctly by learning the query workload. The pruning utilizes a predicted range query that covers the correct neighbors. We show that that our approach has the potential to work effectively in a large real-world trajectory dataset.
Hani Ramadhan, Joonho Kwon
IEEE BigData2
2020 Learning Minimum Bounding Rectangles for Efficient Trajectory Similarity Search
abstract
Early pruning of dissimilar trajectories is important in similar trajectory search on a big mobility data. R-trees can perform the pruning effectively, but the search and index size become inefficient due to numerous overlapping of minimum bounding regions in a dense and big dataset. Thus, we introduce the extended usage of learned index to learn the minimum bounding rectangles for trajectory similarity search. Our approach is designed to provide an effective pruning for trajectory similarity search with less storage size.
Hani Ramadhan, Joonho Kwon
IEEE BigData2
2019 Extracting valid indoor semantic trajectories using movement constraints
abstract
An indoor semantic trajectory is a sequence of timestamped semantic positions inside a building. However, its extraction depends on the erroneous indoor positioning. The error leads to an invalid trajectory that has distant consecutive positions. This invalid trajectory may lead to an issue of the non-sensical patterns when analyzing a big semantic trajectory data. To prevent extracting invalid trajectories, we apply the movement constraints to infer only close positions to the current position. We extend the constraints to several indoor positioning techniques, such as Hidden Markov Model, K-Nearest Neighbor, or Deep Neural Network. We show that our approach can effectively extract valid indoor semantic trajectories.
Hani Ramadhan, Yoga Yustiawan, Joonho Kwon
IEEE BigData3
2015 A timeline visualization system for road traffic big data
abstract
The rapid converging of big data and IoT (Internet of Things) technologies provides more opportunities in the area of road traffic applications. In this paper, we discuss a timeline visualization tool which enables us to better understand of traffic behaviors from road traffic big data.
Ardi Imawan, Joonho Kwon
IEEE BigData2
2015 Scalable extraction of timeline information from road traffic data using MapReduce
abstract
Due to the increasing number of vehicles in recent years, traffic congestion problem is a common issue for residents of metropolises. For a better understanding of traffic congestion, the analyzed data from big data technology can be provided as timeline information. However, a scalability problem would occur when we convert raw traffic data into the timeline information due to the volume and complexity of traffic data. In this paper, we present two MapReduce-based approaches which extract the timeline information of road traffic data. By utilizing the distribute processing strategy, we can resolve the scalability problem. We propose an iterative approach of MapReduce as a baseline approach and a single iteration approach as an efficient solution. The single iteration easily extended to support various and/or complex analytic queries by providing proper codes at a reducer. We validate experimentally our MapReduce-based approaches on real traffic dataset from a Busan Intelligent Transport System (ITS) center.
Ardi Imawan, Fadhilah Kurnia Putri, Seonga An, Han-You Jeong, Joonho Kwon
DSAA5
2012 Non-redundant web services composition based on a two-phase algorithm
Joonho Kwon, Daewook Lee
Data Knowl. Eng.1
2011 A Graph Model Based Simulation Tool for Generating RFID Streaming Data
Haipeng Zhang 0001, Joonho Kwon, Bonghee Hong
APWeb2
2008 Value-based predicate filtering of XML documents
Joonho Kwon, Praveen Rao 0001, Bongki Moon, Sukho Lee
Data Knowl. Eng.1
2005 FiST: Scalable XML Document Filtering by Sequencing Twig Patterns
Joonho Kwon, Praveen Rao 0001, Bongki Moon, Sukho Lee
VLDB1