Kulsawasd Jitkajornwanich

dblp:117/5981 · DBLP profile ↗
← Back
8ranked-venue papers in the field
4as first author
2since 2021 · last 2024
0000-0002-6926-7577ORCID · verified

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 7 (4 first)Database Systems & Data Management · 1
YearPublicationVenuePosition
2024 Leveraging Race Prediction Algorithms to Enhance Team Composition in Big Data Science Teams
abstract
As big data science projects scale in complexity, optimizing team composition has become vital for improving creativity, productivity, and project success. We explore the possibility of incorporating race prediction algorithms for enhancing racial diversity in team composition in big data science projects. This paper evaluates five race prediction algorithms—wru, ethnicolr, ethnicolr2, pyethnicity, and rethnicity—and then discuss their potential in supporting racially diverse team assembly in big data projects. Utilizing three datasets, we assess algorithm performance and applicability, emphasizing their role in building balanced teams that enhance agility, inclusivity, and bias mitigation. We present an actionable methodology for integrating demographic insights into team management. In addition, we propose ethical safeguards to ensure responsible race prediction use, recommending data privacy measures, aggregate-only data handling, and transparency in communication. We argue that when used within ethical constraints, race prediction can support robust team processes, reduce reliance on less diverse teams, and ultimately facilitate more creative and equitable big data project outcomes.
Thanathip Chumthong, Kulsawasd Jitkajornwanich, Obada Kraishan, Kerk F. Kee, Akan Narabin
IEEE Big Data2
2023 A Capacity Framework of Community Readiness for Supporting Big Data Science Projects during Cyberinfrastructure Diffusion
abstract
This study presents a capacity framework for measuring community readiness for supporting big data science projects during cyberinfrastructure (CI) diffusion. CI projects are academic big data science projects driven by data-intensive research efforts. CI projects are an interesting example of understanding big data science projects from a scientific and academic perspective. In this paper, we advanced the argument that in order for CI projects to succeed, they need to draw from five dimensions of community readiness. More specifically, we present a capacity framework of community readiness consisting of the five dimensions of national support networks, peer-to-peer support networks, CI opinion leaders, curriculum and training, and student workforce. We proposed composite scale items developed to quantitatively define and measure these five dimensions, which can be administered using a questionnaire. The overall average score and the composite scores of the five dimensions can be utilized as reflexive assessment and feedback for CI projects about the academic and professional community they belong to. Future research could statistically validate the framework via factor analyses.
Kerk F. Kee, Alex Olshansky, Shan Xu 0001, Kulsawasd Jitkajornwanich
IEEE Big Data4
2018 A Survey of Spatio-Temporal Database Research
Neelabh Pant, Mohammadhani Fouladgar, Ramez Elmasri, Kulsawasd Jitkajornwanich
ACIIDS (2)4
2018 Utilizing Twitter Data for Early Flood Warning in Thailand
abstract
Natural disasters cause significant damage to the country as well as its citizens as we have seen in the news. Drought, wild fire, earthquake and flooding are some examples of the primary natural disasters occurred in Thailand. In this research, we focus on "flooding" and use data from Twitter, where users' mobile devices are utilized as IoT input channels. The goal of this work is to analyze near real-time data (tweets) for early flood warning. Traditional methods in processing, analyzing and reporting a flooding event take quite some time. In social medias (through cellphones), on the other hand, by harvesting crowdsources, potential flooding can be predicted faster-though with the price of reliability of the retrieved tweets. In our research, several techniques are incorporated in order to maximize the accuracy of results, including, tokenization, geo-encoding and decoding, NLP via string matching (Levenshtein's algorithms), and Google APIs for visualization. Finally, the dynamic yet user-friendly map is produced with respect to the posted relevant tweets along their associated frequencies.
Kulsawasd Jitkajornwanich, Chanwit Kongthong, Nattaya Khongsoontornjaroen, Jeedapa Kaiyasuan, Siam Lawawirojwong, Panu Srestasathiern, Siwapon Srisonphan, Peerapon Vateekul
IEEE BigData1
2017 Road map extraction from satellite imagery using connected component analysis and landscape metrics
abstract
Road map extraction is considered an essential task in GIS as its results are the basis of location-based applications in various domains. Examples include GPS navigation on cell phone, delivery route optimization and planning, tourist attraction locator, and location-based marketing. Satellite imagery, one of the big spatial data sources, is used in this research - though other types of remotely-sensed images can also be applied, such as aerial photographs from aircrafts, UAVs or drones. Despite several methods and techniques proposed and accompanied with different performance criteria, the focus was mainly on the accuracy aspect rather than the completeness of the result sets. That is, the results were said to be satisfactory if it met a certain accuracy criterion associated with some benchmark data sets, regardless of whether all results were retrieved. In many cases; however, both accuracy and completeness are equally important. In this paper, we enhance the result accuracy by incorporating connected component analysis into the method as well as the completeness performance by utilizing an ecology concept, called Landscape Metrics, which describes spatial characteristics, patterns, and correlations of areas/patches through different indices. Two types of metrics are used: shape metrics and isolation metrics. The performance is evaluated based on four criteria: precision, recall, quality, and F1 scores. The results show that more than 90% of performance is achieved in all four criteria.
Kulsawasd Jitkajornwanich, Peerapon Vateekul, Teerapong Panboonyuen, Siam Lawawirojwong, Siwapon Srisonphan
IEEE BigData1
2017 Ocean surface current prediction based on HF radar observations using trajectory-oriented association rule mining
abstract
HF (high frequency) coastal radar system is used to capture the surface current behavior - in terms of velocity and direction - in the ocean near the coast. 18 HF coastal radar stations were implemented along the Gulf of Thailand in order to monitor for disasters (e.g., Tsunami) as well as relevant risks. The HF systems are also to serve other life-critical applications, such as water quality control and monitoring, chemical spill backtracking, and marine navigation. However, not all the applications can benefit from this near-real-time HF data; some applications in different domains require forecast values. The examples include search-and-rescue system and hazardous materials spill trajectory prediction. Therefore, in this paper, we propose a predictive model for future current data based on historical HF coastal radar data sets, utilizing association rule mining combined with an object dispersion concept. So, the full potential of HF radar systems can be exploited. The spatial and temporal dimensions are taken into account when designing our predictive system, which consists of two phases: ocean surface current track formulation and spatio-temporal association rule mining. The experiments are performed on a two-year HF radar dataset (2014-2015) using Google Cloud Platform. The resulting forecast current values: velocity and direction are then compared with testing datasets (using 10-fold cross validation) of the actual recorded values and evaluated based on percentage accuracy and RMSE, respectively.
Kulsawasd Jitkajornwanich, Peerapon Vateekul, Upa Gupta, Teeranai Kormongkolkul, Arnon Jirakittayakorn, Siam Lawawirojwong, Siwapon Srisonphan
IEEE BigData1
2016 Adapting K-means clustering to identify spatial patterns in storms
abstract
This paper extends our previous work on deriving meaningful storm patterns from very large rainfall data. In an earlier work, we described MapReduce-based algorithms to identify three types of the storms: local, hourly and overall storms. In general, local storms have temporal characteristics of the storms at a particular site, hourly storms have spatial characteristics of the storms at a particular hour and overall storms have both spatial and temporal characteristics of the storm. We aim to find meaningful patterns and predict trajectories in the spatio-temporal data (i.e. overall storms which are sets of geographically overlapping, consecutive hourly storms). In this paper, we adapt K-Means clustering to find different types of hourly storms based on their shapes and sizes. Since the rainfall data are typically larger than the memory capacity of a single computer, we have implemented this clustering algorithm in Apache Spark, which is a distributed data processing framework, and have run our experiments on a computer cluster.
Upa Gupta, Kulsawasd Jitkajornwanich, Ramez Elmasri, Leonidas Fegaras
IEEE BigData2
2013 Complete storm identification algorithms from big raw rainfall data using MapReduce framework
abstract
In our previous work, we described various aspects of our approach in converting big raw rainfall data into meaningful storm concepts. Three concepts were defined: local, hourly, and overall storms. The latter describes overall spatio-temporal characteristics of a storm as it progresses over time. We previously described MapReduce-based algorithms for local and hourly storm identification. Overall storms are the most complex to identify, and are at the core of the storm identification system. Multiple consecutive hourly storms that have spatial overlap are combined to create storm-centric characteristics of the whole storm, which could not be captured in most existing hydrology research. In this paper, we propose a MapReduce-based overall storm identification algorithm, which is based on iteration on MapReduce framework. This greatly improves performance when compared to the depth-first search (DFS) graph traveling approach as introduced in our previous work. In addition, additional essential storm characteristics of hourly and overall storms are introduced in this paper. Examples include storm center concepts for hourly storms and storm track and speed for overall storms.
Kulsawasd Jitkajornwanich, Upa Gupta, Sakthi Kumaran Shanmuganathan, Ramez Elmasri, Leonidas Fegaras, John McEnery
IEEE BigData1