Hong-Gee Kim

dblp:75/4867 · DBLP profile ↗
← Back
52ranked-venue papers
1as first author
23since 2021 · last 2026
0000-0002-2610-4321ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 1 first-author · 17 since 2021Databases, data management, data science and information retrieval · 15 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 3Systems, architecture and hardware · 2Theory of computation · 1
YearPublicationVenuePosition
2026 CL²GEC: A Multi-Discipline Benchmark for Continual Learning in Chinese Literature Grammatical Error Correction
abstract
Shang Qin, Jingheng Ye, Yinghui Li, Hai-Tao Zheng, Qi Li, Jinxiao Shan, Zhixing Li, Hong-Gee Kim. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Shang Qin, Jingheng Ye, Hai-Tao Zheng 0002, Qi Li 0002, Jinxiao Shan, Hong-Gee Kim
ACL (1)8
2026 GMSA: Enhancing Context Compression via Group Merging and Layer Semantic Alignment
abstract
Jiwei Tang, Zhicheng Zhang, Shunlong Wu, Jingheng Ye, Lichen Bai, Zitai Wang, Tingwei Lu, Lin Hai, Yiming Zhao, Hai-Tao Zheng, Hong-Gee Kim. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jiwei Tang, Zhicheng Zhang 0008, Shunlong Wu, Jingheng Ye, Lichen Bai, Zitai Wang, Tingwei Lu, Lin Hai, Hai-Tao Zheng 0002, Hong-Gee Kim
ACL (1)11
2026 Incremental extraction of bespoke association rules
Eung-Hee Kim, Hong-Gee Kim, Suk-hyung Hwang
Knowl. Based Syst.2
2026 Generalized few-shot intent detection by prompt learning without forgetting
Chaiyut Luoyiching, Yangning Li, Rongsheng Li, Zhixiong Cao, Hai-Tao Zheng 0002, Hanjing Su, Hong-Gee Kim
Neural Comput. Appl.9
2026 Synonym Knowledge Graph Enhanced Language Model for Inconsistent Hallucination Detection
abstract
Neural sequence models, despite their proficiency in generating highly fluent sentences, have also exhibited a tendency to hallucinate, introducing additional content that lacks grounding in the input data, as evidenced by recent investigations. This variety of fluent yet erroneous outputs poses a significant challenge, as it is difficult for users to discern the veracity of the presented content and identify inaccuracies. Several methods have been proposed to address this challenge. However, they sometimes detect the synonym of the true token in the generation as hallucination. To alleviate the above problem, we propose a novel architecture, i.e., Synonym Knowledge Graph Enhanced Language Model (SKGELM). We construct and prune a synonym knowledge graph according to the source sentence, which can help the subsequently graph attention network and classifier detect hallucination. Empirical results on SUMMAC benchmark and multi-domain Chinese–English translation benchmark show that our method achieves state-of-the-art performance and improves the best baseline significantly.
Hai-Tao Zheng 0002, Hui Wang 0030, Hong-Gee Kim
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2025 Frozen Language Models Are Gradient Coherence Rectifiers in Vision Transformers
abstract
Large language models (LLMs) have demonstrated remarkable performance in multimodal tasks even with frozen LLM Block and only a few trainable parameters. However, the underlying mechanisms of how LLMs enhance multimodal performance remains unclear. In this work, we focus on the phenomenon that ``Merely concatenating a frozen LLM block to the Vision Transformer (ViT) encoder can yield significant performance enhancements. Moreover, the choice of LLM block and insertion position can have a substantial impact, leading to varying degrees of improvement''. We analyze the optimization of the training process from the perspective of gradient dynamics and find that frozen LLM blocks act as gradient coherence rectifiers, aligning the gradients of different samples more closely during training. Furthermore, we demonstrate that the representation similarity between the inserted LLM block and the adjacent ViT block influences performance, with greater similarity tending to yield larger positive gains. Through these findings, we can justify the selection of suitable LLM blocks to be inserted at appropriate positions, and introduce additional gradient backpropagation paths by incorporating LLM blocks, could improve the performance of vanilla ViT through the rectification effect of gradient consistency during the training process, without the need to add LLM blocks during inference. Our experiments demonstrate the effectiveness of this strategy, making the practical application of the gradient rectification effect feasible.
Lichen Bai, Zixuan Xiong, Xiangjin Xie, Ruijie Guo, Zhanhui Kang, Hai-Tao Zheng 0002, Hong-Gee Kim
AAAI9
2025 CLEME2.0: Towards Interpretable Evaluation by Disentangling Edits for Grammatical Error Correction
abstract
Jingheng Ye, Zishan Xu, Yinghui Li, Linlin Song, Qingyu Zhou, Hai-Tao Zheng, Ying Shen, Wenhao Jiang, Hong-Gee Kim, Ruitong Liu, Xin Su, Zifei Shan. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Jingheng Ye, Zishan Xu, Linlin Song, Qingyu Zhou, Hai-Tao Zheng 0002, Ying Shen 0001, Hong-Gee Kim, Zifei Shan
ACL (1)9
2025 Dual Denoising Diffusion Model for Session-based Social Recommendation
abstract
Session-based Social Recommendation (SSR) enhances item recommendations by incorporating both session interactions and social network data. Despite recent progress, existing SSR methods-primarily based on Graph Neural Networks-are highly susceptible to session noise (irrelevant or unintentional interactions) and social noise (misleading signals from connected users). Prior denoising strategies often rely on heuristic resampling or reweighting techniques, which lack generalizability and robustness across diverse datasets. In this work, we explore a novel direction by introducing diffusion models for denoising in SSR. However, applying diffusion to SSR presents unique challenges due to heterogeneous data modalities, incompatible noise patterns, and the absence of semantic guidance during the reverse process. To overcome these challenges, we propose D3MRec, a Dual Denoising Diffusion Model specifically designed for SSR. D3MRec employs a dual-branch architecture that independently models session sequences and social graphs, applying denoising diffusion in their respective hidden representation spaces. This decoupled design preserves the structural integrity of each modality while enabling modality-specific denoising. Moreover, we introduce cross-modal guidance by leveraging collaborative signals from the other branch during the reverse diffusion process, enhancing alignment between session intents and social preferences. The dual denoising processes not only mitigate noise within each modality but also serve as mutual priors, facilitating robust and consistent representation learning across modalities. Extensive experiments on multiple benchmarks show that D3MRec significantly outperforms state-of-the-art models, particularly under noisy conditions, demonstrating its effectiveness and robustness.
Mengying Lu, Hai-Tao Zheng 0002, Qi Li 0002, Jinxiao Shan, Hong-Gee Kim
CIKM7
2025 LexSemBridge: Fine-Grained Dense Representation Enhancement Through Token-Aware Embedding Augmentation
abstract
As queries in retrieval-augmented generation (RAG) pipelines powered by large language models (LLMs) become increasingly complex and diverse, dense retrieval models have demonstrated strong performance in semantic matching. Nevertheless, they often struggle with fine-grained retrieval tasks, where precise keyword alignment and span-level localization are required, even in cases with high lexical overlap that would intuitively suggest easier retrieval. To systematically evaluate this limitation, we introduce two targeted tasks, keyword retrieval and part-of-passage retrieval, designed to simulate practical fine-grained scenarios. Motivated by these observations, we propose LexSemBridge, a unified framework that enhances dense query representations through fine-grained, input-aware vector modulation. LexSemBridge constructs latent enhancement vectors from input tokens using three paradigms: Statistical (SLR), Learned (LLR), and Contextual (CLR), and integrates them with dense embeddings via element-wise interaction. Theoretically, we show that this modulation preserves the semantic direction while selectively amplifying discriminative dimensions. LexSemBridge operates as a plug-in without modifying the backbone encoder and naturally extends to both text and vision modalities. Extensive experiments across semantic and fine-grained retrieval tasks validate the effectiveness and generality of our approach. All code and models are publicly available at https://github.com/Jasaxion/LexSemBridge/
Shaoxiong Zhan, Hongming Tan, Xiaodong Cai, Hai-Tao Zheng 0002, Zifei Shan, Hong-Gee Kim
ECAI9
2025 Exploring the Implicit Semantic Ability of Multimodal Large Language Models: A Pilot Study on Entity Set Expansion
abstract
The rapid development of multimodal large language models (MLLMs) has brought significant improvements to a wide range of tasks in realworld applications. However, LLMs still exhibit certain limitations in extracting implicit semantic information. In this paper, we applies MLLMs to the Multi-modal Entity Set Expansion (MESE) task, which aims to expand a handful of seed entities with new entities belonging to the same semantic class, and multi-modal information is provided with each entity. We explore the capabilities of MLLMs to understand implicit semantic information at the entity-level granularity through the MESE task, introducing a listwise ranking method LUSAR that maps local scores to global rankings. Our LUSAR demonstrates significant improvements in MLLM’s performance on the MESE task, marking the first use of generative MLLM for ESE tasks and extending the applicability of listwise ranking.
Hebin Wang, Yangning Li, Hai-Tao Zheng 0002, Hong-Gee Kim
ICASSP6
2025 Efficient Visual Storytelling through Descriptive Words Distillation and Dynamic Decoding
abstract
Visual storytelling, a complex task in natural language generation, aims to create coherent and engaging narratives from a sequence of images, requiring more intricate and lengthy descriptions than typical image captioning. Current methods generally employ sophisticated modal interaction modules and require substantial additional data for training. This paper introduces two innovative techniques to advance visual storytelling capabilities. First, we propose a method to distill descriptive word embeddings from model parameters as an extension to the model’s pre-trained embeddings. This approach focuses the training process on these descriptive words, allowing for efficient fine-tuning and eliminates the need for additional visual modules and reduces computational overhead. Second, we develop a dynamic decoding strategy that adapts the text generation process based on the characteristic of the image features. This ensures a improved alignment with image features and the generated stories. Our experiments on the VIST dataset demonstrate that these methods achieve an impressive performance. Ablation studies show significant improvements in training efficiency and narrative quality compared to existing approaches.
Zixuan Xiong, Lichen Bai, Hai-Tao Zheng 0002, Hong-Gee Kim
ICASSP6
2025 SmellDetector: Multi-Label Code Smell Detection and Refactoring with Large Language Models
abstract
Large Language Models (LLMs) have demonstrated impressive capabilities in many tasks such as code generation and automated program repair. However, code LLMs have ignored another important task in programmers’ daily development work, which is to improve the maintainability, readability, and scalability of the program. All of these characteristics are related to code smells and we study how to improve them by detecting and removing code smells. Most works on code smells still rely on using measures formulated by experts as features, but lack of use of the rich prior knowledge contained in code LLMs. In this paper, we propose SmellDetector, a comprehensive model for both code smell detection and refactoring opportunities detection in Java. We train the model with the designed prompt which contains both code smells of class-level and method-level in the same code snippet, including more than 20 types. We achieve state-of-the-art performance on the code smell detection task and change the basic paradigm of code smell detection from binary classification problem to multi-label classification. Finally, it has been verified through experiments that good code smell detection helps to detect refactoring opportunities.
Hai-Tao Zheng 0002, Haiye Lin, Hong-Gee Kim, Bingxu An, Zhao Wei, Yong Xu 0007
IJCNN6
2025 Repository-Level Code Smell Detection Based on Multi-Scale Code Information and LLM Assistance
abstract
In software engineering, code smell detection has always been an important research task because it affects the readability and maintainability of programs. Traditional code smell detection work mostly focuses on a single file, and the study of repo-level code smell caused by the interaction of multiple files is still relatively lacking, although it is more practical in reality. Although large language models (LLMs) have recently achieved remarkable success in the field of code generation, simply copying the fine-tuning method based on LLM is not good enough in this task because it is difficult to model the logical relationship between multiple files. In this paper, we propose a repo-level code smell detection model(RSD), which divides the levels according to the distance between the rest class and the problem class, uses cross attention and CNN to model global and local information, and uses LLM to generate teacher vectors in the same scenario for assistance. Finally, the detection performance surpasses the fine-tuning baseline method based on multiple code LLMs and has achieved the state-of-the-art. Finally, we also contribute the BenchMark used in this article to help researchers in the repo-level smell detection task.
Yongqin Zeng, Hai-Tao Zheng 0002, Haiye Lin, Hong-Gee Kim, Bingxu An, Zhao Wei, Yong Xu 0007
IJCNN6
2025 Eliminating Retrieval Knowledge Conflicts: Cross-Validation Re-ranking with Large Language Models
abstract
In retrieval-augmented generation (RAG) systems, Large Language Models (LLMs) have been shown to be effective for re-ranking. However, existing research often prioritizes passage relevance over reliability, which can result in the incorporation of conflicting information and the generation of ambiguous responses. This issue becomes particularly pronounced when addressing inter-context knowledge conflicts, where candidate documents present contradictory information that may mislead the model. To mitigate this problem, we propose a novel cross-validation re-ranking technique designed specifically to resolve inter-context knowledge conflicts during the retrieval process. We also develop a new dataset, ContraPRT, to evaluate the ability of models to rank passages containing conflicting knowledge. Experimental results using GPT-4 and LlaMA3-70B demonstrate that our approach not only effectively filters out conflicting information but also ensures accurate passage rankings, thereby providing reliable supplementary knowledge for the generation module.
Qirui Wu, Lin Hai, Hai-Tao Zheng 0002, Ruobing Xie, Saiyong Yang, Xingwu Sun, Zhanhui Kang, Hong-Gee Kim
IJCNN8
2025 Alleviating Chinese repetitive generation via intra and intersentence penalty
Fangqing Jiang, Hui Wang 0030, Hai-Tao Zheng 0002, Hong-Gee Kim
Neural Comput. Appl.5
2024 Depth Aware Hierarchical Replay Continual Learning for Knowledge Based Question Answering
abstract
Continual learning is an emerging area of machine learning that deals with the issue where models adapt well to the latest data but lose the ability to remember past data due to changes in the data source. A widely adopted solution is by keeping a small memory of previous learned data that use replay. Most of the previous studies on continual learning focused on classification tasks, such as image classification and text classification, where the model needs only to categorize the input data. Inspired by the human ability to incrementally learn knowledge and solve different problems using learned knowledge, we considered a more pratical scenario, knowledge based quesiton answering about continual learning. In this scenario, each single question is different from others(means different fact trippes to answer them) while classification tasks only need to find feature boundaries of different categories, which are the curves or surfaces that separate different categories in the feature space. To address this issue, we proposed a depth aware hierarchical replay framework which include a tree structure classfier to have a sense of knowledge distribution and fill the gap between text classfication tasks and question-answering tasks for continual learning, a local sampler to grasp these critical samples and a depth aware learning network to reconstructe the feature space of a single learning round. In our experiments, we have demonstrated that our proposed model outperforms previous continual learning methods in mitigating the issue of catastrophic forgetting.
Zhixiong Cao, Hai-Tao Zheng 0002, Yangning Li, Rongsheng Li, Hong-Gee Kim
LREC/COLING6
2024 Multi-label classification with XGBoost for metabolic pathway prediction
abstract
BACKGROUND: Metabolic pathway prediction is one possible approach to address the problem in system biology of reconstructing an organism's metabolic network from its genome sequence. Recently there have been developments in machine learning-based pathway prediction methods that conclude that machine learning-based approaches are similar in performance to the most used method, PathoLogic which is a rule-based method. One issue is that previous studies evaluated PathoLogic without taxonomic pruning which decreases its performance. RESULTS: In this study, we update the evaluation results from previous studies to demonstrate that PathoLogic with taxonomic pruning outperforms previous machine learning-based approaches and that further improvements in performance need to be made for them to be competitive. Furthermore, we introduce mlXGPR, a XGBoost-based metabolic pathway prediction method based on the multi-label classification pathway prediction framework introduced from mlLGPR. We also improve on this multi-label framework by utilizing correlations between labels using classifier chains. We propose a ranking method that determines the order of the chain so that lower performing classifiers are placed later in the chain to utilize the correlations between labels more. We evaluate mlXGPR with and without classifier chains on single-organism and multi-organism benchmarks. Our results indicate that mlXGPR outperform other previous pathway prediction methods including PathoLogic with taxonomic pruning in terms of hamming loss, precision and F1 score on single organism benchmarks. CONCLUSIONS: The results from our study indicate that the performance of machine learning-based pathway prediction methods can be substantially improved and can even outperform PathoLogic with taxonomic pruning.
Hyunwhan Joe, Hong-Gee Kim
BMC Bioinform.2
2024 A Segment Augmentation and Prediction Consistency Framework for Multi-label Unknown Intent Detection
abstract
Multi-label unknown intent detection is a challenging task where each utterance may contain not only multiple known but also unknown intents. To tackle this challenge, pioneers proposed to predict the intent number of the utterance first, then compare it with the results of known intent matching to decide whether the utterence contains unknown intent(s). Though they have made remarkable progress on this task, their methods still suffer from two important issues: (1) It is inadequate to extract multiple intents using only utterance encoding; (2) Optimizing two sub-tasks (intent number prediction and known intent matching) independently leads to inconsistent predictions. In this article, we propose to incorporate segment augmentation rather than only use utterance encoding to better detect multiple intents. We also design a prediction consistency module to bridge the gap between the two sub-tasks. Empirical results on MultiWOZ2.3 and MixSNIPS datasets show that our method achieves state-of-the-art performance and significantly improves the best baseline.
Miaoxin Chen, Cao Liu, Boqi Dai, Hai-Tao Zheng 0002, Hui Wang 0030, Rui Xie 0005, Hong-Gee Kim
ACM Trans. Knowl. Discov. Data8
2023 Attention Gate Between Capsules in Fully Capsule-Network Speech Recognition
Kyungmin Lee, Hyeontaek Lim, Munhwan Lee, Hong-Gee Kim
INTERSPEECH4
2022 Modeling Latent Autocorrelation for Session-based Recommendation
abstract
Session-based Recommendation (SBR) aims to predict the next item for the current session, which consists of several clicked items in a short period by an anonymous user. Most of the sequential modeling approaches to SBR are focusing on adopting advanced Deep Neural Networks (DNNs), and these methods require increasingly longer training times. Existing studies have shown that some traditional SBR methods can outperform some DNN-based sequential models, however, few studies have attempted to investigate the effectiveness of traditional methods in recent years. In this paper, we propose a novel and concise SBR model inspired by the basic concept of autocorrelation in the Stochastic Process. Autocorrelation measures the correlation of a process at different moments. Therefore, it is natural to use it to model the correlation of clicked item sequences at different time shifts. Specifically, we use Fast Fourier Transforms (FFT) to compute the autocorrelation and combine it with several linear transformations to enhance the session representation. By this means, our proposed method can learn better session preferences and is more efficient than most DNN-based models. Extensive experiments on two public datasets show that the proposed method outperforms state-of-the-art models in both effectiveness and efficiency.
Xianghong Xu 0001, Kai Ouyang, Liuyin Wang, Jiaxin Zou, Yanxiong Lu, Hai-Tao Zheng 0002, Hong-Gee Kim
CIKM7
2022 Diversify Search Results Through Graph Attentive Document Interaction
Xianghong Xu 0001, Kai Ouyang, Yanxiong Lu, Hai-Tao Zheng 0002, Hong-Gee Kim
DASFAA (1)6
2022 A novel approach to predicting the synergy of anti-cancer drug combinations using document-based feature extraction
abstract
BACKGROUND: To reduce drug side effects and enhance their therapeutic effect compared with single drugs, drug combination research, combining two or more drugs, is highly important. Conducting in-vivo and in-vitro experiments on a vast number of drug combinations incurs astronomical time and cost. To reduce the number of combinations, researchers classify whether drug combinations are synergistic through in-silico methods. Since unstructured data, such as biomedical documents, include experimental types, methods, and results, it can be beneficial extracting features from documents to predict anti-cancer drug combination synergy. However, few studies predict anti-cancer drug combination synergy using document-extracted features. RESULTS: We present a novel approach for anti-cancer drug combination synergy prediction using document-based feature extraction. Our approach is divided into two steps. First, we extracted documents containing validated anti-cancer drug combinations and cell lines. Drug and cell line synonyms in the extracted documents were converted into representative words, and the documents were preprocessed by tokenization, lemmatization, and stopword removal. Second, the drug and cell line features were extracted from the preprocessed documents, and training data were constructed by feature concatenation. A prediction model based on deep and machine learning was created using the training data. The use of our features yielded higher results compared to the majority of published studies. CONCLUSIONS: Using our prediction model, researchers can save time and cost on new anti-cancer drug combination discoveries. Additionally, since our feature extraction method does not require structuring of unstructured data, new data can be immediately applied without any data scalability issues.
Yongsun Shim, Munhwan Lee, Pil-Jong Kim, Hong-Gee Kim
BMC Bioinform.4
2021 Sequential routing framework: Fully capsule network-based speech recognition
Kyungmin Lee, Hyunwhan Joe, Hyeontaek Lim, Kwangyoun Kim, Chang Woo Han, Hong-Gee Kim
Comput. Speech Lang.7
2017 Constructing faceted taxonomy for heterogeneous entities based on object properties in linked data
Nansu Zong, Hong-Gee Kim, Sejin Nam
Data Knowl. Eng.2
2017 xStore: Federated temporal query processing for large scale RDF triples on a cloud environment
Jae-Hong Eom, Sejin Nam, Nansu Zong, Dong-Hyuk Im, Hong-Gee Kim
Neurocomputing6
2017 A dynamic and parallel approach for repetitive prime labeling of XML with MapReduce
Dong-Hyuk Im, Taewhi Lee, Hong-Gee Kim
J. Supercomput.4
2015 FARM: An FCA-based Association Rule Miner
Eung-Hee Kim, Hong-Gee Kim, Suk-hyung Hwang, Sungin Lee
Knowl. Based Syst.2
2015 Aligning ontologies with subsumption and equivalence relations in Linked Data
Nansu Zong, Sejin Nam, Jae-Hong Eom, Hyunwhan Joe, Hong-Gee Kim
Knowl. Based Syst.6
2015 SigMR: MapReduce-based SPARQL query processing by signature encoding and multi-way join
Dong-Hyuk Im, Hong-Gee Kim
J. Supercomput.3
2014 STEP: An ontology-based smart clinical document template editing and production system
Sejin Nam, Sungin Lee, James G. Boram Kim, Hong-Gee Kim
Expert Syst. Appl.4
2014 Exploiting social bookmarking services to build clustered user interest profile for personalized search
Sungin Lee, Hong-Gee Kim
Inf. Sci.3
2012 Shared decision support system on dental restoration
Seon Gyu Park, Sungin Lee, Myeng-Ki Kim, Hong-Gee Kim
Expert Syst. Appl.4
2011 Using Folksonomies for Building User Interest Profile
Hong-Gee Kim
UMAP2
2011 OntoPipeliner: An ontology-based automatic semantic service pipeline generator
Sungin Lee, Senator Jeong, Hong-Gee Kim, Hanmin Jung, Mikyoung Lee, Seung-Jae Song, Beom-Jong You
Expert Syst. Appl.3
2011 Mining and Representing User Interests: The Case of Tagging Practices
abstract
Social tagging in online communities has become an important method for reflecting classified thoughts of individual users. A number of social Web sites provide tagging functionalities and also offer folksonomies within or across the sites. However, it is practically not easy to find users' interests based on such folksonomies. In this paper, we provide a novel approach for clustering user-centric interests by analyzing tagging practices of individual users. To do this, we collect Really Simple Syndication data from blogosphere, find conceptual clusters using formal concept analysis, and then evaluate the significance of these clusters. The results of the empirical evaluation show that we can effectively recommend different collections of tags to an individual or a set of users.
Hak Lae Kim, John G. Breslin, Stefan Decker, Hong-Gee Kim
IEEE Trans. Syst. Man Cybern. Part A4
2010 The Use of Ontology in Dental Restorative Treatment Decision Support System
abstract
Finding appropriate caries treatments is of paramount importance in dental decision making. That is, finding restorative treatment alternatives predicated on the dental disease and findings are advantageous and gainful in dental restorative decision making. The most immediate problem in clinical decision support systems in dentistry is to capture a doctor's clinical knowledge of treatments. This study is to specify the inter-relations among disease, anatomy, and treatment for restorative treatment decision support, and to conceptualize restorative treatment. As an explanatory example, we expound the developmental process of our ontology, and the formal approach used, for caries treatment.
Seon Gyu Park, Sungin Lee, Myeng-Ki Kim, Hong-Gee Kim
FOIS4
2010 Mitigation of Large-scale RDF Data Loading with the Employment of a Cloud Computing Service
Hyun Namgoong, Hong-Gee Kim
KEOD3
2010 Improving the Workflow of Semantic Web Portals using M/R in Cloud Platform
Seokchan Yun, Mina Song, Hyun Namgung, Sung-Kwon Yang, Hong-Gee Kim
KEOD6
2010 GOClonto: An ontological clustering approach for conceptualizing PubMed abstracts
Hai-Tao Zheng 0002, Charles Borchert, Hong-Gee Kim
J. Biomed. Informatics3
2009 Browsing Unbounded Social, Linked Data Instances with User Perspectives
abstract
With the emerging uses of semantically enriched social data on the Web, linked data are expected to envision a next generation of the current web. As dasiaweb of datapsila, they are spread as pieces of data into the Web with links to related objects or concepts. Data instances distributed with URIs, those that enable identification and combination of data instances, can be consumed with shared data vocabularies. Easy and intuitive access to the data should be provided for data-centered uses of the Web. This paper introduces a linked data browser providing an intuitive view, especially helping casual userspsila understanding of data instances and their relationships. The browser also satisfies the requirement of a generic browser: handling unexpected domains of data across the links. By adapting userpsilas perspectives captured during browsing, the browser enables users to view any types of linked data instances with different views pertinent to their intentions and types of data.
Hyun Namgoong, Sung-Kwon Yang, Mina Song, Hong-Gee Kim
ASONAM4
2009 Exploiting noun phrases and semantic relationships for text document clustering
Hai-Tao Zheng 0002, Bo-Yeong Kang, Hong-Gee Kim
Inf. Sci.3
2009 Are you an invited speaker? A bibliometric analysis of elite groups for scholarly events in bioinformatics
abstract
Abstract Participating in scholarly events (e.g., conferences, workshops, etc.) as an elite‐group member such as an organizing committee chair or member, program committee chair or member, session chair, invited speaker, or award winner is beneficial to a researcher's career development. The objective of this study is to investigate whether elite‐group membership for scholarly events is representative of scholars' prominence, and which elite group is the most prestigious. We collected data about 15 global (excluding regional) bioinformatics scholarly events held in 2007. We sampled (via stratified random sampling) participants from elite groups in each event. Then, bibliometric indicators (total citations and h index) of seven elite groups and a non‐elite group, consisting of authors who submitted at least one paper to an event but were not included in any elite group, were observed using the Scopus Citation Tracker. The Kruskal–Wallis test was performed to examine the differences among the eight groups. Multiple comparison tests (Dwass, Steel, Critchlow–Fligner) were conducted as follow‐up procedures. The experimental results reveal that scholars in an elite group have better performance in bibliometric indicators than do others. Among the elite groups, the invited speaker group has statistically significantly the best performance while the other elite‐group types are not significantly distinguishable. From this analysis, we confirm that elite‐group membership in scholarly events, at least in the field of bioinformatics, can be utilized as an alternative marker for a scholar's prominence, with invited speaker being the most important prominence indicator among the elite groups.
Senator Jeong, Sungin Lee, Hong-Gee Kim
J. Assoc. Inf. Sci. Technol.3
2009 Exploiting corpus-related ontologies for conceptualizing document corpora
abstract
Abstract As a greater volume of information becomes increasingly available across all disciplines, many approaches, such as document clustering and information visualization, have been proposed to help users manage information easily. However, most of these methods do not directly extract key concepts and their semantic relationships from document corpora, which could help better illuminate the conceptual structures within given information. To address this issue, we propose an approach called “Clonto” to process a document corpus, identify the key concepts, and automatically generate ontologies based on these concepts for the purpose of conceptualization. For a given document corpus, Clonto applies latent semantic analysis to identify key concepts, allocates documents based on these concepts, and utilizes WordNet to automatically generate a corpus‐related ontology. The documents are linked to the ontology through the key concepts. Based on two test collections, the experimental results show that Clonto is able to identify key concepts, and outperforms four other clustering algorithms. Moreover, the ontologies generated by Clonto show significant informative conceptual structures.
Hai-Tao Zheng 0002, Charles Borchert, Hong-Gee Kim
J. Assoc. Inf. Sci. Technol.3
2009 Two-Phase Chief Complaint Mapping to the UMLS Metathesaurus in Korean Electronic Medical Records
abstract
The task of automatically determining the concepts referred to in chief complaint (CC) data from electronic medical records (EMRs) is an essential component of many EMR applications aimed at biosurveillance for disease outbreaks. Previous approaches that have been used for this concept mapping have mainly relied on term-level matching, whereby the medical terms in the raw text and their synonyms are matched with concepts in a terminology database. These previous approaches, however, have shortcomings that limit their efficacy in CC concept mapping, where the concepts for CC data are often represented by associative terms rather than by synonyms. Therefore, herein we propose a concept mapping scheme based on a two-phase matching approach, especially for application to Korean CCs, which uses term-level complete matching in the first phase and concept-level matching based on concept learning in the second phase. The proposed concept-level matching suggests the method to learn all the terms (associative terms as well as synonyms) that represent the concept and predict the most probable concept for a CC based on the learned terms. Experiments on 1204 CCs extracted from 15,618 discharge summaries of Korean EMRs showed that the proposed method gave significantly improved F-measure values compared to the baseline system, with improvements of up to 73.57%.
Bo-Yeong Kang, Daewon Kim 0001, Hong-Gee Kim
IEEE Trans. Inf. Technol. Biomed.3
2008 The State of the Art in Tag Ontologies: A Semantic Model for Tagging and Folksonomies
Hak Lae Kim, Simon Scerri, John G. Breslin, Stefan Decker, Hong-Gee Kim
Dublin Core Conference5
2008 Social Semantic Cloud of Tag: Semantic Model for Social Tagging
Hak Lae Kim, John G. Breslin, Sung-Kwon Yang, Hong-Gee Kim
KES-AMSTA4
2008 A Concept-Driven Automatic Ontology Generation Approach for Conceptualization of Document Corpora
abstract
In the age of increasing information availability, many techniques, such as document clustering and information visualization, have been developed to ease understanding of information for users. However, most of these methods do not help users directly understand key concepts and their semantic relationships in document corpora, which are critical for capturing their conceptual structures. Therefore, we propose a novel approach called 'Clonto' to identify the key concepts and automatically generate ontologies based on these concepts for conceptualization of document corpora. Clonto applies latent semantic analysis to identify key concepts, allocates documents based on these concepts, and utilizes WordNet to automatically generate a corpus-related ontology. The documents are linked to the ontology through the key concepts. The experimental results show that Clonto can identify key concepts with a high precision and the clustering results of Clonto outperform the STC (Suffix Tree Clustering) algorithm, the Lingo clustering algorithm, the Fuzzy Ants clustering algorithm, and clustering based on TRS (Tolerance Rough Set). Moreover, based on the same document corpus, the ontology generated by Clonto shows a significant informative conceptual structure.
Hai-Tao Zheng 0002, Charles Borchert, Hong-Gee Kim
Web Intelligence3
2008 An ontology-based approach to learnable focused crawling
Hai-Tao Zheng 0002, Bo-Yeong Kang, Hong-Gee Kim
Inf. Sci.3
2007 WANT: A Personal Knowledge Management System on Social Software Agent Technologies
Hak Lae Kim, Jaehwa Choi, Hong-Gee Kim, Suk-hyung Hwang
KES-AMSTA3
2007 An Agent Environment for Contextualizing Folksonomies in a Triadic Context
Hong-Gee Kim, Suk-hyung Hwang, Yu-Kyung Kang, Hak Lae Kim, Hae Sool Yang
KES-AMSTA1
2006 A Data-Driven Approach to Constructing an Ontological Concept Hierarchy Based on the Formal Concept Analysis
Suk-hyung Hwang, Hong-Gee Kim, Myeng-Ki Kim, Sung-Hee Choi, Hae Sool Yang
ICCSA (4)2
2005 A FCA-Based Ontology Construction for the Design of Class Hierarchy
Suk-hyung Hwang, Hong-Gee Kim, Hae Sool Yang
ICCSA (3)2