Kangil Kim

dblp:45/8372 · DBLP profile ↗
← Back
24ranked-venue papers
8as first author
9since 2021 · last 2026
0000-0003-3220-6401ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 8 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 What and when to look? Temporal span proposal network for video relation detection
Sangmin Woo, Junhyug Noh, Kangil Kim
Expert Syst. Appl.3
2025 RSCF: Relation-Semantics Consistent Filter for Entity Embedding of Knowledge Graph
abstract
In knowledge graph embedding, leveraging relation specific entity transformation has markedly enhanced performance.However, the consistency of embedding differences before and after transformation remains unaddressed, risking the loss of valuable inductive bias inherent in the embeddings.This inconsistency stems from two problems.First, transformation representations are specified for relations in a disconnected manner, allowing dissimilar transformations and corresponding entity embeddings for similar relations.Second, a generalized plug-in approach as a SFBR (Semantic Filter Based on Relations) disrupts this consistency through excessive concentration of entity embeddings under entity-based regularization, generating indistinguishable score distributions among relations.In this paper, we introduce a plug-in KGE method, Relation-Semantics Consistent Filter (RSCF).Its entity transformation has three features for enhancing semantic consistency: 1) shared affine transformation of relation embeddings across all relations, 2) rooted entity transformation that adds an entity embedding to its change represented by the transformed vector, and 3) normalization of the change to prevent scale reduction.To amplify the advantages of consistency that preserve semantics on embeddings, RSCF adds relation transformation and prediction modules for enhancing the semantics.In knowledge graph completion tasks with distance-based and tensor decomposition models, RSCF significantly outperforms state-of-the-art KGE methods, showing robustness across all relations and their frequencies.
Jinwook Park, Kangil Kim
ACL (1)3
2025 Probability Distribution Collapse: A Critical Bottleneck to Compact Unsupervised Neural Grammar Induction
abstract
Unsupervised neural grammar induction aims to learn interpretable hierarchical structures from language data.However, existing models face an expressiveness bottleneck, often resulting in unnecessarily large yet underperforming grammars.We identify a core issue, probability distribution collapse, as the underlying cause of this limitation.We analyze when and how the collapse emerges across key components of neural parameterization and introduce a targeted solution, collapse-relaxing neural parameterization, to mitigate it.Our approach substantially improves parsing performance while enabling the use of significantly more compact grammars across a wide range of languages, as demonstrated through extensive empirical analysis.
Jinwook Park, Kangil Kim
EMNLP2
2025 Preference Distillation via Value based Reinforcement Learning
abstract
Direct Preference Optimization (DPO) is a powerful paradigm to align language models with human preferences using pairwise comparisons. However, its binary win-or-loss supervision often proves insufficient for training small models with limited capacity. Prior works attempt to distill information from large teacher models using behavior cloning or KL divergence. These methods often focus on mimicking current behavior and overlook distilling reward modeling. To address this issue, we propose \textit{Teacher Value-based Knowledge Distillation} (TVKD), which introduces an auxiliary reward from the value function of the teacher model to provide a soft guide. This auxiliary reward is formulated to satisfy potential-based reward shaping, ensuring that the global reward structure and optimal policy of DPO are preserved. TVKD can be integrated into the standard DPO training framework and does not require additional rollouts. Our experimental results show that TVKD consistently improves performance across various benchmarks and model sizes.
Minchan Kwon, Junwon Ko, Kangil Kim, Junmo Kim 0002
NeurIPS3
2024 Label-Focused Inductive Bias over Latent Object Features in Visual Classification
abstract
Most neural networks for classification primarily learn features differentiated by input-domain related information such as visual similarity of objects in an image. While this focus is natural behavior, it can inadvertently introduce an inductive bias that conflicts with unseen relations in an implicit output-domain determined by human labeling based on their own world knowledge. Such conflicts can limit generalization of models by potential dominance of the input-domain focused bias in inference. To overcome this limitation without external resources, we introduce Output-Domain focused Biasing (ODB) training strategy that constructs inductive biases on features differentiated by only output labels. It has four steps: 1) it learns intermediate latent object features in an unsupervised manner; 2) it decouples their visual dependencies by assigning new independent embedding parameters; 3) it captures structured features optimized for the original classification task; and 4) it integrates the structured features with the original visual features for the final prediction. We implement the ODB on a vision transformer architecture, and achieved significant improvements on image classification benchmarks. This paper offers a straightforward and effective method to obtain and utilize output-domain focused inductive bias for classification mapping two different domains.
Ilmin Kang, HyounYoung Bae, Kangil Kim
ICLR3
2024 Fixed Non-negative Orthogonal Classifier: Inducing Zero-mean Neural Collapse with Feature Dimension Separation
abstract
Fixed classifiers in neural networks for classification problems have demonstrated cost efficiency and even outperformed learnable classifiers in some popular benchmarks when incorporating orthogonality. Despite these advantages, prior research has yet to investigate the training dynamics of fixed orthogonal classifiers on neural collapse, a recently clarified phenomenon that last-layer features converge to a specific form, called simplex ETF, in training classification models involving the post-zero-error phase. Ensuring this phenomenon is critical for obtaining global optimality in a layer-peeled model, potentially leading to enhanced performance in practice. However, fixed orthogonal classifiers cannot invoke neural collapse due to their geometric limitations. To overcome the limits, we analyze a $\textit{zero-mean neural collapse}$ considering the orthogonality in non-negative Euclidean space. Then, we propose a $\textit{fixed non-negative orthogonal classifier}$ that induces the optimal solution and maximizes the margin of an orthogonal layer-peeled model by satisfying the properties of zero-mean neural collapse. Building on this foundation, we exploit a $\textit{feature dimension separation}$ effect inherent in our classifier for further purposes: (1) enhances softmax masking by mitigating feature interference in continual learning and (2) tackles the limitations of mixup on the hypersphere in imbalanced learning. We conducted comprehensive experiments on various datasets and architectures and demonstrated significant performance improvements.
Hoyong Kim, Kangil Kim
ICLR2
2023 Feature structure distillation with Centered Kernel Alignment in BERT transferring
Hee-Jun Jung, Seung-Hoon Na, Kangil Kim
Expert Syst. Appl.4
2023 Tackling the Challenges in Scene Graph Generation With Local-to-Global Interactions
abstract
In this work, we seek new insights into the underlying challenges of the scene graph generation (SGG) task. Quantitative and qualitative analysis of the visual genome (VG) dataset implies: 1) ambiguity: even if interobject relationship contains the same object (or predicate), they may not be visually or semantically similar; 2) asymmetry: despite the nature of the relationship that embodied the direction, it was not well addressed in previous studies; and 3) higher-order contexts: leveraging the identities of certain graph elements can help generate accurate scene graphs. Motivated by the analysis, we design a novel SGG framework, Local-to-global interaction networks (LOGINs). Locally, interactions extract the essence between three instances of subject, object, and background, while baking direction awareness into the network by explicitly constraining the input order of subject and object. Globally, interactions encode the contexts between every graph component (i.e., nodes and edges). Finally, Attract and Repel loss is utilized to fine-tune the distribution of predicate embeddings. By design, our framework enables predicting the scene graph in a bottom-up manner, leveraging the possible complementariness. To quantify how much LOGIN is aware of relational direction, a new diagnostic task called Bidirectional Relationship Classification (BRC) is also proposed. Experimental results demonstrate that LOGIN can successfully distinguish relational direction than existing methods (in BRC task), while showing state-of-the-art results on the VG benchmark (in SGG task).
Sangmin Woo, Junhyug Noh, Kangil Kim
IEEE Trans. Neural Networks Learn. Syst.3
2022 Spherization Layer: Representation Using Only Angles
abstract
In neural network literature, angular similarity between feature vectors is frequently used for interpreting or re-using learned representations. However, the inner product in neural networks partially disperses information over the scales and angles of the involved input vectors and weight vectors. Therefore, when using only angular similarity on representations trained with the inner product, information loss occurs in downstream methods, which limits their performance. In this paper, we proposed the $\textit{spherization layer}$ to represent all information on angular similarity. The layer 1) maps the pre-activations of input vectors into the specific range of angles, 2) converts the angular coordinates of the vectors to Cartesian coordinates with an additional dimension, and 3) trains decision boundaries from hyperplanes, without bias parameters, passing through the origin. This approach guarantees that representation learning always occurs on the hyperspherical surface without the loss of any information unlike other projection-based methods. Furthermore, this method can be applied to any network by replacing an existing layer. We validate the functional correctness of the proposed method in a toy task, retention ability in well-known image classification tasks, and effectiveness in word analogy test and few-shot learning. Code is publicly available at https://github.com/GIST-IRR/spherization_layer
Hoyong Kim, Kangil Kim
NeurIPS2
2020 Effect of pre-training to build a regression model using shallow neural network for semiconductor plasma etch process equipment
abstract
Plasma etch process is one of manufacturing steps to fabricate semiconductor chips and the difficulty of etch process has become harder and harder because the target specifications of chips have become harsher. To overcome this circumstance, monitoring plasma parameters in real-time has been requested. In this study, regression models to predict a plasma density from the intensities of optical wavelength obtained from plasma etch process chamber using shallow neural network were presented. The estimation results of several models with or without pre-training were also analyzed. The model using variational auto-encoder showed the best performance and it can expect to be easily accepted in semiconductor industry because optical intensity measurement device was already equipped for plasma etch process chamber.
Ohyung Kwon, Nayeon Lee, Kangil Kim
IEEE BigData3
2019 Improving LSTM CRFs using character-based compositions for Korean named entity recognition
Seung-Hoon Na, Hyun Kim 0003, Jinwoo Min, Kangil Kim
Comput. Speech Lang.4
2019 Transition-Based Korean Dependency Parsing Using Hybrid Word Representations of Syllables and Morphemes with LSTMs
abstract
Recently, neural approaches for transition-based dependency parsing have become one of the state-of-the art methods for performing dependency parsing tasks in many languages. In neural transition-based parsing, a parser state representation is first computed from the configuration of a stack and a buffer, which is then fed into a feed-forward neural network model that predicts the next transition action. Given that words are basic elements of a stack and buffer, a parser state representation is considerably affected by how a word representation is defined. In particular, word representation issues become more critical in morphologically rich languages such as Korean, as the set of potential words is not bound but introduce the second-order vocabulary complexity, called the phrase vocabulary complexity due to the agglutinative characteristics of the language. In this article, we propose a hybrid word representation that combines two compositional word representations, each of which is derived from representations of syllables and morphemes , respectively. Our underlying assumption for this hybrid word representation is that, because both syllables and morphemes are two common ways of decomposing Korean words, it is expected that their effects in inducing word representation are complementary to one another. Experimental results carried on Sejong and SPMRL 2014 datasets show that our proposed hybrid word representation leads to the state-of-the-art performance.
Seung-Hoon Na, Jianri Li, Jong-Hun Shin, Kangil Kim
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2018 Verbosity normalized pseudo-relevance feedback in information retrieval
Seung-Hoon Na, Kangil Kim
Inf. Process. Manag.2
2017 Time-Sensitive Adaptation of Regularization Strength of Recurrent Neural Networks for Accurate Learning
abstract
Regularization is an important issue for neural networks because of strong expression power causing overfitting to data. A regularization method is to penalize cost functions by activation-based penalty. In its applications to recurrent neural networks, the method usually assigns penalty uniformly distributed over time steps. However, required strength for recurrent networks differs by time steps. In this paper we propose a new activation-based penalty function varying its strength over time steps in recurrent neural networks. To verify its impact, we conducted practical experiments to predict the power consumption of home appliances. In the results, the proposed method reduced training errors and maintained validation and test errors, which implies the improvement of forecasting ability. In sensitivity analysis, the method restricted sudden decrease of impact of early time steps to the cost.
Kangil Kim
ICMLA1
2017 Divergence-based fine pruning of phrase-based statistical translation model
Kangil Kim, Eun-Jin Park, Jong-Hun Shin, Oh-Woog Kwon, Young Kil Kim
Comput. Speech Lang.1
2017 Center-shared sliding ensemble of neural networks for syntax analysis of natural language
Kangil Kim, Yun Jin, Seung-Hoon Na, Young Kil Kim
Expert Syst. Appl.1
2016 Recursion-Based Biases in Stochastic Grammar Model Genetic Programming
abstract
The estimation of distribution algorithms (EDAs) applied to genetic programming (GP) have been studied by a number of authors. Like all EDAs, they suffer from biases induced by the model building and sampling process. However, the biases are amplified in the algorithms for GP. In particular, many systems use stochastic grammars as their model representation, but biases arise due to grammar recursion. We define and estimate the bias due to recursion in grammar-based EDAs in GP, using methods derived from computational linguistics. We confirm the extent of bias in some simple experimental examples. We then propose some methods to repair this bias. We apply the estimation of bias, and its repair, to some more practical applications. We experimentally demonstrate the extent of bias arising from recursion, and the performance improvements that can result from correcting it.
Kangil Kim, Robert I. McKay, Nguyen Xuan Hoai
IEEE Trans. Evol. Comput.1
2013 Cutting evaluation costs: An investigation into early termination in genetic programming
abstract
Genetic programming is very computationally intensive, particularly in CPU time. A number of approaches to evaluation cost reduction have been proposed, among them early termination of evaluation (applicable in problem domains where estimates of the final fitness value are available during evaluation). Like all cost reduction techniques, early termination balances overall computation cost against the risk of finding worse solutions. We evaluate the influence of various properties of the problem domain - problem class, reliability of fitness estimates, trajectory of fitness estimates, and evolutionary trajectory - to determine whether any is able to predict the effects of early termination. There is little correlation with any of these, with one exception. Boolean problems see little change in running time, and hence only small changes in performance, are distinguished by both problem class, and each of the other metrics.
Namyong Park 0001, Kangil Kim, Robert I. McKay
IEEE Congress on Evolutionary Computation2
2013 Stochastic Diversity Loss and Scalability in Estimation of Distribution Genetic Programming
abstract
In estimation of distribution algorithms (EDAs), probability models hold accumulating evidence on the location of an optimum. Stochastic sampling drift has been heavily researched in EDA optimization but not in EDAs applied to genetic programming (EDA-GP). We show that, for EDA-GPs using probabilistic prototype tree models, stochastic drift in sampling and selection is a serious problem, inhibiting scaling to complex problems. Problems requiring deep dependence in their probability structure see such rapid stochastic drift that the usual methods for controlling drift are unable to compensate. We propose a new alternative, analogous to likelihood weighting of evidence. We demonstrate in a small-scale experiment that it does counteract the drift, sufficiently to leave EDA-GP systems subject to similar levels of stochastic drift to other EDAs.
Kangil Kim, Robert I. McKay
IEEE Trans. Evol. Comput.1
2012 Implicit bias and recursive grammar structures in estimation of distribution genetic programming
abstract
Much recent research in Estimation of Distribution Algorithms (EDA) applied to Genetic Programming has adopted a Stochastic Context Free Grammar(SCFG)-based model formalism. However these methods generate biases which may be indistinguishable from selection bias, resulting in sub-optimal performance. The primary factor generating this bias is the combined effect of recursion in the grammars and depth limitation removing some sample trees from the distribution. Here, we demonstrate the bias and provide exact estimates of its scale (assuming infinite populations and simple recursions). We define a quantity h which determines both whether bias occurs (h >; 1) and its scale. We apply this analysis to a number of simple illustrative grammars, and to a range of practically-used GP grammars, showing that this bias is both real and important.
Kangil Kim, Nguyen Xuan Hoai, Robert I. McKay
IEEE Congress on Evolutionary Computation1
2012 Analysing the Effects of Diverse Operators in a Genetic Programming System
Minhyeok Kim 0001, Robert I. McKay, Kangil Kim, Nguyen Xuan Hoai
PPSN (1)3
2011 Operator Self-adaptation in Genetic Programming
Minhyeok Kim 0001, Robert I. McKay, Nguyen Xuan Hoai, Kangil Kim
EuroGP4
2011 Structural difficulty in estimation of distribution genetic programming
abstract
Estimation of Distribution Algorithms were introduced into Genetic Programming over 15 years ago, and have demonstrated good performance on a range of problems, but there has been little research into their limitations. We apply two such algorithms - scalar and vectorial Stochastic Grammar GP - to Daida's well-known Lid problem, to better understand their ability to learn specific structures. The scalar algorithm performs poorly, but the vectorial version shows good overall performance. We then extended Daida's problem to explore the vectorial algorithm's ability to find even more specific structures, finding that the performance fell off rapidly as the specificity of the required structure increased. Thus although this particular system has less severe structural difficulty issues than standard GP, it is by no means free of them. Track: Genetic Programming
Kangil Kim, Minhyeok Kim 0001, Robert I. McKay
GECCO1
2010 Sampling Bias in Estimation of Distribution Algorithms for Genetic Programming Using Prototype Trees
Kangil Kim, Robert I. McKay, Dharani Punithan
PRICAI1