Inwon Kang

dblp:70/3770 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 9 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Language Model Representations for Efficient Few-Shot Tabular Classification
abstract
The Web is a rich source of structured data in the form of tables, from product catalogs and knowledge bases to scientific datasets. However, the heterogeneity of the structure and semantics of these tables makes it challenging to build a unified method that can effectively leverage the information they contain. Meanwhile, Large language models (LLMs) are becoming an increasingly integral component of web infrastructure for tasks like semantic search. This raises a crucial question: can we leverage these already-deployed LLMs to classify structured data in web-native tables (e.g., product catalogs, knowledge base exports, scientific data portals), avoiding the need for specialized models or extensive retraining? This work investigates a lightweight paradigm, $\textbf{Ta}$ble $\textbf{R}$epresentation with $\textbf{L}$anguage Model~($\textbf{TaRL}$), for few-shot tabular classification that directly utilizes semantic embeddings of individual table rows. We first show that naive application of these embeddings underperforms compared to specialized tabular models. We then demonstrate that their potentials can be unlocked with two key techniques: removing the common component from all embeddings and calibrating the softmax temperature. We show that a simple meta-learner, trained on handcrafted features, can learn to predict an appropriate temperature. This approach achieves performance comparable to state-of-the-art models in low-data regimes ($k \leq 32$) of semantically-rich tables. Our findings demonstrate the viability of reusing existing LLM infrastructure for efficient semantics-driven pathway to reuse existing LLM infrastructure for Web table understanding.
Inwon Kang, Parikshit Ram, Yi Zhou 0015, Horst Samulowitz, Oshani Seneviratne
WWW1
2024 Effective Data Distillation for Tabular Datasets (Student Abstract)
abstract
Data distillation is a technique of reducing a large dataset into a smaller dataset. The smaller dataset can then be used to train a model which can perform comparably to a model trained on the full dataset. Past works have examined this approach for image datasets, focusing on neural networks as target models. However, tabular datasets pose new challenges not seen in images. A sample in tabular dataset is a one dimensional vector unlike the two (or three) dimensional pixel grid of images, and Non-NN models such as XGBoost can often outperform neural network (NN) based models. Our contribution in this work is two-fold: 1) We show in our work that data distillation methods from images do not translate directly to tabular data; 2) We propose a new distillation method that consistently outperforms the baseline for multiple different models, including non-NN models such as XGBoost.
Inwon Kang, Parikshit Ram, Yi Zhou 0015, Horst Samulowitz, Oshani Seneviratne
AAAI1
2024 Enhancing Web Spam Detection Through a Blockchain-Enabled Crowdsourcing Mechanism
Noah Kader, Inwon Kang, Oshani Seneviratne
WISE (5)2
2022 Landslide Likelihood Prediction using Machine Learning Algorithms
abstract
The supply of electricity via power plants is critical to the operation of many critical infrastructure systems in modern society. Natural hazards can disrupt the power supply, cause power outages that can halt economic growth, and impede emergency response until power is restored. The proposed work aims to predict the landslides likelihood in these critical infrastructure locations in the Northeastern USA using integrated databases of explanatory variables and machine learning algorithms. First, data related to landslides are obtained and merged, including topographic, soil moisture, and precipitation-related data. Five regression algorithms, namely: Random Forest, Extreme Gradient Boosting (XGBoost), K-Nearest Neighbor regression (KNN), Linear Support Vector Regressor (SVR), and Linear regression, are utilized to predict the landslide probability and evaluated on the dataset. The accuracy of the models is assessed by using statistical metrics such as mean absolute error (MAE), mean squared error (MSE), and root mean squared error (RMSE). The study results show that Random Forest outperformed other models with the mutual information feature selection method. It achieved an MSE of 0.0011 with mutual information-based feature selection and an MSE of 0.00157 without feature selection. KNN regressor outperformed the other models with an MSE of 0.00139 with correlation-based information selection. The proposed landslide identification model with Random Forest algorithm shows outstanding robustness and great potential in tackling the landslide likelihood prediction by employing ML algorithms.
Vasundhara Acharya, Anindita Ghosh, Inwon Kang, Thilanka Munasinghe, K. C. Binita
IEEE Big Data3
2022 Blockchain Interoperability Landscape
abstract
Blockchain has become a popular emergent technology in many industries. It is suitable for a broad range of applications, from its base role as an immutable distributed ledger to the deployment of distributed applications. Many organizations are adopting the technology, but choosing a specific blockchain implementation in an emerging field exposes them to significant technology risk. Selecting the wrong implementation could expose an organization to security vulnerabilities, reduce access to its target audience, or cause issues in the future when switching to a more mature protocol. Blockchain interoperability aims to solve this adaptability problem by increasing the extensibility of blockchain, enabling the addition of new use cases and features without sacrificing the performance of the original blockchain. However, most existing blockchain platforms need to be designed for interoperability, and simple operations like sending assets across platforms create problems. Cryptographic protocols that are secure in isolation may become insecure when several different (individually secure) protocols are composed. Similarly, utilizing trusted custodians may undercut most of the benefits of decentralization offered by blockchain-based systems. Even though there is some research and development in the field of blockchain interoperability, a characterization of the interoperability solutions for various infrastructure options is lacking. This paper presents a methodology for characterizing blockchain interoperability solutions that will help focus on new developments and evaluate existing and future solutions in this space.
Inwon Kang, Aparna Gupta, Oshani Seneviratne
IEEE Big Data1
2022 Crowdsourcing Perceptions of Gerrymandering
abstract
Gerrymandering is the manipulation of redistricting to influence the results of a set of elections for local representatives. Gerrymandering has the potential to drastically swing power in legislative bodies even with no change in a population’s political views. Identifying gerrymandering and measuring fairness using metrics of proposed district plans is a topic of current research, but there is less work on how such plans will be perceived by voters. Gathering data on such perceptions presents several challenges such as the ambiguous definitions of ‘fair’ and the complexity of real world geography and district plans. We present a dataset collected from an online crowdsourcing platform on a survey asking respondents to mark which of two maps of equal population distribution but different districts appear more ‘fair’ and the reasoning for their decision. We performed preliminary analysis on this data and identified which of several commonly suggested metrics are most predictive of the responses. We found that the maximum perimeter of any district was the most predictive metric, especially with participants who reported that they made their decision based on the shape of the districts.
Benjamin Kelly, Inwon Kang, Lirong Xia
HCOMP2
2008 Understanding individual investor's behavior with financial information disclosed on the web sites
abstract
Today's financial firms are required to disclose a great deal of data—including investment support information—on their corporate web sites. Since web sites have become an integral part of financial information disclosures, and because the multimedia characteristics embedded in those web sites have been shown to affect investors' responses, our research examines the various factors influencing investors' intention to use financial web sites to search for information. The basic premise of this study is that the reactions of individual investors in such situations translate naturally to an intention to use financial web sites and, ultimately, to actual use of these sites. By using a technology acceptance model, we conducted a rigorous questionnaire survey, over an illustrative web site on which financial information is disclosed on a regular basis as a means of providing individual investors with decision support. Our principal findings showed that: (1) consistency and technical convenience influence perceived ease of use; (2) decision quality, investment information and information quality affect perceived usefulness; and (3) perceived usefulness to the individual investor is affected most by decision quality, while perceived ease of use is influenced equally by consistency and technical convenience.
Kun Chang Lee, Namho Chung, Inwon Kang
Behav. Inf. Technol.3
2006 Identifying Critical Dependencies Using General Bayesian Network: An Application to Aviation Operation
abstract
This paper proposes the general Bayesian network as a tool for identifying critical dependencies among complex human beliefs, especially for the case of airline service industry. When we consider an overall efficacy of the flight operation, to understand underlying beliefs between pilots and controllers is critical. However, even for the experts in organizational behavior, it is difficult to identify dependencies among such the beliefs. To tackle this ambiguous issue, we suggest the use of the general Bayesian network to automatically generate interesting hypotheses. The experimental results reveal the possibility of selective attention to the critical dependencies for achieving enhanced efficacy.
Sung Woo Shin, Inwon Kang
SNPD2
2006 Erratum to "Efficiency analysis of controls in EDI applications" [Information & Management 42 (2005) 425-439]
Sangjae Lee, Kidong Lee, Inwon Kang
Inf. Manag.3
2005 Efficiency analysis of controls in EDI applications
Sangjae Lee, Kidong Lee, Inwon Kang
Inf. Manag.3
2005 KMPI: measuring knowledge management performance
Kun Chang Lee, Sangjae Lee, Inwon Kang
Inf. Manag.3
2005 Erratum to "KMPI: Measuring knowledge management performance" [Information & Management 42 (2005) 469-482]
Kun Chang Lee, Sangjae Lee, Inwon Kang
Inf. Manag.3
2004 Using fuzzy cognitive map for the relationship management in airline service
Inwon Kang, Sangjae Lee, Jiho Choi
Expert Syst. Appl.1