Inwon Kang

dblp:70/3770 · DBLP profile ↗
← Back
9ranked-venue papers in the field
2as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 4Information Retrieval & Web Search · 2 (1 first)Big Data, Cloud & Distributed Data Systems · 2 (1 first)Data Mining & Knowledge Discovery · 1
YearPublicationVenuePosition
2026 Language Model Representations for Efficient Few-Shot Tabular Classification
abstract
The Web is a rich source of structured data in the form of tables, from product catalogs and knowledge bases to scientific datasets. However, the heterogeneity of the structure and semantics of these tables makes it challenging to build a unified method that can effectively leverage the information they contain. Meanwhile, Large language models (LLMs) are becoming an increasingly integral component of web infrastructure for tasks like semantic search. This raises a crucial question: can we leverage these already-deployed LLMs to classify structured data in web-native tables (e.g., product catalogs, knowledge base exports, scientific data portals), avoiding the need for specialized models or extensive retraining? This work investigates a lightweight paradigm, $\textbf{Ta}$ble $\textbf{R}$epresentation with $\textbf{L}$anguage Model~($\textbf{TaRL}$), for few-shot tabular classification that directly utilizes semantic embeddings of individual table rows. We first show that naive application of these embeddings underperforms compared to specialized tabular models. We then demonstrate that their potentials can be unlocked with two key techniques: removing the common component from all embeddings and calibrating the softmax temperature. We show that a simple meta-learner, trained on handcrafted features, can learn to predict an appropriate temperature. This approach achieves performance comparable to state-of-the-art models in low-data regimes ($k \leq 32$) of semantically-rich tables. Our findings demonstrate the viability of reusing existing LLM infrastructure for efficient semantics-driven pathway to reuse existing LLM infrastructure for Web table understanding.
Inwon Kang, Parikshit Ram, Yi Zhou 0015, Horst Samulowitz, Oshani Seneviratne
WWW1
2024 Enhancing Web Spam Detection Through a Blockchain-Enabled Crowdsourcing Mechanism
Noah Kader, Inwon Kang, Oshani Seneviratne
WISE (5)2
2022 Landslide Likelihood Prediction using Machine Learning Algorithms
abstract
The supply of electricity via power plants is critical to the operation of many critical infrastructure systems in modern society. Natural hazards can disrupt the power supply, cause power outages that can halt economic growth, and impede emergency response until power is restored. The proposed work aims to predict the landslides likelihood in these critical infrastructure locations in the Northeastern USA using integrated databases of explanatory variables and machine learning algorithms. First, data related to landslides are obtained and merged, including topographic, soil moisture, and precipitation-related data. Five regression algorithms, namely: Random Forest, Extreme Gradient Boosting (XGBoost), K-Nearest Neighbor regression (KNN), Linear Support Vector Regressor (SVR), and Linear regression, are utilized to predict the landslide probability and evaluated on the dataset. The accuracy of the models is assessed by using statistical metrics such as mean absolute error (MAE), mean squared error (MSE), and root mean squared error (RMSE). The study results show that Random Forest outperformed other models with the mutual information feature selection method. It achieved an MSE of 0.0011 with mutual information-based feature selection and an MSE of 0.00157 without feature selection. KNN regressor outperformed the other models with an MSE of 0.00139 with correlation-based information selection. The proposed landslide identification model with Random Forest algorithm shows outstanding robustness and great potential in tackling the landslide likelihood prediction by employing ML algorithms.
Vasundhara Acharya, Anindita Ghosh, Inwon Kang, Thilanka Munasinghe, K. C. Binita
IEEE Big Data3
2022 Blockchain Interoperability Landscape
abstract
Blockchain has become a popular emergent technology in many industries. It is suitable for a broad range of applications, from its base role as an immutable distributed ledger to the deployment of distributed applications. Many organizations are adopting the technology, but choosing a specific blockchain implementation in an emerging field exposes them to significant technology risk. Selecting the wrong implementation could expose an organization to security vulnerabilities, reduce access to its target audience, or cause issues in the future when switching to a more mature protocol. Blockchain interoperability aims to solve this adaptability problem by increasing the extensibility of blockchain, enabling the addition of new use cases and features without sacrificing the performance of the original blockchain. However, most existing blockchain platforms need to be designed for interoperability, and simple operations like sending assets across platforms create problems. Cryptographic protocols that are secure in isolation may become insecure when several different (individually secure) protocols are composed. Similarly, utilizing trusted custodians may undercut most of the benefits of decentralization offered by blockchain-based systems. Even though there is some research and development in the field of blockchain interoperability, a characterization of the interoperability solutions for various infrastructure options is lacking. This paper presents a methodology for characterizing blockchain interoperability solutions that will help focus on new developments and evaluate existing and future solutions in this space.
Inwon Kang, Aparna Gupta, Oshani Seneviratne
IEEE Big Data1
2022 Crowdsourcing Perceptions of Gerrymandering
abstract
Gerrymandering is the manipulation of redistricting to influence the results of a set of elections for local representatives. Gerrymandering has the potential to drastically swing power in legislative bodies even with no change in a population’s political views. Identifying gerrymandering and measuring fairness using metrics of proposed district plans is a topic of current research, but there is less work on how such plans will be perceived by voters. Gathering data on such perceptions presents several challenges such as the ambiguous definitions of ‘fair’ and the complexity of real world geography and district plans. We present a dataset collected from an online crowdsourcing platform on a survey asking respondents to mark which of two maps of equal population distribution but different districts appear more ‘fair’ and the reasoning for their decision. We performed preliminary analysis on this data and identified which of several commonly suggested metrics are most predictive of the responses. We found that the maximum perimeter of any district was the most predictive metric, especially with participants who reported that they made their decision based on the shape of the districts.
Benjamin Kelly, Inwon Kang, Lirong Xia
HCOMP2
2006 Erratum to "Efficiency analysis of controls in EDI applications" [Information & Management 42 (2005) 425-439]
Sangjae Lee, Kidong Lee, Inwon Kang
Inf. Manag.3
2005 Efficiency analysis of controls in EDI applications
Sangjae Lee, Kidong Lee, Inwon Kang
Inf. Manag.3
2005 KMPI: measuring knowledge management performance
Kun Chang Lee, Sangjae Lee, Inwon Kang
Inf. Manag.3
2005 Erratum to "KMPI: Measuring knowledge management performance" [Information & Management 42 (2005) 469-482]
Kun Chang Lee, Sangjae Lee, Inwon Kang
Inf. Manag.3