Donghee Han 0001

dblp:196/1720-1 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0002-4560-6810ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Every Preference Has Its Strength: Injecting Ordinal Semantics into LLM-Based Recommenders
abstract
Recent work has shown that large language models (LLMs) can enhance recommender systems by integrating collaborative filtering (CF) signals through hybrid prompting. However, most existing CF-LLM frameworks collapse explicit ratings into implicit or positive-only feedback, discarding the ordinal structure that conveys fine-grained preference strength. As a result, these models struggle to exploit graded semantics and nuanced preference distinctions. We propose Ordinal Semantic Anchoring (OSA), a hybrid CF-LLM framework that explicitly incorporates preference strength by modeling interaction-level user feedback. OSA represents ordinal preference levels as numeric textual tokens and uses their token embeddings as semantic anchors to align user-item interaction representations in the LLM latent space. Through strength-aware alignment across ordinal levels, OSA preserves preference semantics when integrating collaborative signals with LLMs. Experiments on multiple real-world datasets demonstrate that OSA consistently outperforms existing baselines, particularly in pairwise preference evaluation, highlighting its effectiveness in modeling fine-grained user preferences over prior CF-LLM methods.
Jiwon Jeong, Donghee Han 0001, Sungrae Hong, Woosung Kang 0001, Mun Yong Yi
SIGIR2
2025 RAG-based Unanswerable Question Detection in Clinical Text-to-SQL
abstract
Large-scale language models (LLMs) have shown exceptional performance in various tasks, particularly in zero-shot and few-shot settings. However, in sensitive domains like healthcare, detecting unanswerable questions remains a critical challenge. This task is challenging due to data imbalance, and existing methods are computationally expensive and inflexible to data distribution changes. To address these issues, we propose Retrieval-augmented Question Answerability Detection (RaQAD), a training-free method that uses LLMs to identify unanswerable questions by retrieving semantically similar examples as few-shot prompts. RaQAD ensures semantically similar sampling, adapts to schema changes, and eliminates the need for additional training. Extensive experiments on clinical datasets demonstrate its effectiveness in outperforming existing approaches while addressing data imbalance challenges.
Donghee Han 0001, Seungjae Lim, Mun Yong Yi
CIKM1
2025 Leveraging LLM-Generated Schema Descriptions for Unanswerable Question Detection in Clinical Data
abstract
Recent advancements in large language models (LLMs) have boosted research on generating SQL queries from domain-specific questions, particularly in the medical domain. A key challenge is detecting and filtering unanswerable questions. Existing methods often relying on model uncertainty, but these require extra resources and lack interpretability. We propose a lightweight model that predicts relevant database schemas to detect unanswerable questions, enhancing interpretability and addressing the data imbalance in binary classification tasks. Furthermore, we found that LLM-generated schema descriptions can significantly enhance the prediction accuracy. Our method provides a resource-efficient solution for unanswerable question detection in domain-specific question answering systems.
Donghee Han 0001, Seungjae Lim, Daeyoung Roh, Sangryul Kim, Sehyun Kim, Mun Yong Yi
COLING1
2025 CTIP: Towards Accurate Tabular-to-Image Generation for Tire Footprint Generation
abstract
Generating images directly from tabular data while ensuring an accurate representation of ground truth is a useful application in manufacturing. Simply embedding tabular data to use it as a condition in image generation models often fails to learn the correspondence between tabular features and their impact on the generated image. To overcome this limitation, we propose Contrastive Tabular-Image Pre-training (CTIP), inspired by the CLIP framework. These pre-train methods help improve the quality of the embedding of the tabular encoder on the tabular data, which then helps improve the performance of the image generation model. CTIP uses contrastive learning on multiple tabular and image data pairs, allowing the model to learn how changes in certain tabular features affect images. This approach is particularly crucial in manufacturing, where accurate capture of product outcomes under varying conditions is essential. We demonstrate that applying CTIP enhances image generation performance, yielding images that closely match ground truth images, even in Feature Few-shot or Feature Zero-shot scenarios where specific features are sparse or novel. We further show the application of CTIP in tire development, where tire footprint images are generated based on tire specifications and test conditions. CTIP produces high-quality embeddings that align well with ground truth images and effectively handle the scarcity or sparseness of specific features, addressing common challenges in new product development. Our code is available in https://github.com/Noverse0/CTIP.git.
Daeyoung Roh, Donghee Han 0001, Jihyun Nam, Jungsoo Oh, Youngbin You, Jeongheon Park, Mun Yong Yi
WACV2
2025 Fine-grained multi-prompt essay scoring with multi-level disentanglement
abstract
Abstract The application of language models in essay scoring has gained significant attention in recent years, typically on the basis of evaluating a single model across multiple prompts. However, in a multi-prompt setup, it is crucial to understand the varying aspects of different prompts. In such settings, there exist notable variations even in a trait with the same name across prompts. This semantic variation on the same traits underscores the need to treat them differently at a fine-grained level according to each prompt. In this study, we propose a multi-level disentanglement framework for multi-prompt essay scoring, designed to achieve fine-grained disentanglement of semantic differences across such traits. Our method not only improves the quality of the essay scoring, but also reduces memory usage and latency. Experimental results highlight that our framework surpasses seven state-of-the-art essay scoring methods and large language model(LLM)-based zero-shot and few-shot approaches, achieving the highest agreement with human essay ratings.
Donghee Han 0001, Daeyoung Roh, Euihwan Han, Hwanjun Song, Mun Yong Yi
Data Min. Knowl. Discov.1
2024 Closer through commonality: Enhancing hypergraph contrastive learning with shared groups
abstract
Hypergraphs provide a superior modeling frame-work for representing complex multidimensional relationships in the context of real-world interactions that often occur in groups, overcoming the limitations of traditional homogeneous graphs. However, there have been few studies on hypergraph-based contrastive learning, and existing graph-based contrastive learning methods have not been able to fully exploit the high-order correlation information in hypergraphs. Here, we propose a Hypergraph Fine-grained contrastive learning (HyFi) method designed to exploit the complex high-dimensional information inherent in hypergraphs. While avoiding traditional graph augmentation methods that corrupt the hypergraph topology, the proposed method provides a simple and efficient learning augmentation function by adding noise to node features. Furthermore, we expands beyond the traditional dichotomous relationship between positive and negative samples in contrastive learning by introducing a new relationship of weak positives. It demonstrates the importance of fine-graining positive samples in contrastive learning. Therefore, HyFi is able to produce high-quality embeddings, and outperforms both supervised and unsupervised baselines in average rank on node classification across 10 datasets. Our approach effectively exploits high-dimensional hypergraph information, shows significant improvement over existing graph-based contrastive learning methods, and is efficient in terms of training speed and GPU memory cost. The source code is available at https://github.com/Noverse0/HyFi.git.
Daeyoung Roh, Donghee Han 0001, Keejun Han, Mun Yong Yi
IEEE Big Data2
2024 Keyword-enhanced recommender system based on inductive graph matrix completion
abstract
Going beyond the user–item rating information, recent studies have utilized additional information to improve the performance of recommender systems. Graph neural network (GNN) based approaches are among the most common. However, existing models that utilize text data require a lot of computing resources and have a complex structure that makes them difficult to utilize in real-world applications. In this research, we propose a new method, keyword-enhanced graph matrix completion (KGMC), which utilizes keyword sharing relationships in user–item graphs. Our model has a simpler structure and requires less computing resources than existing models that utilize text data, but it has the advantage of cross-domain transferability while providing an intuitive understanding of the inference results. KGMC consists of three steps: (1) keyword extraction from the review text, (2) subgraph extraction and keyword-enhanced subgraph construction, and (3) GNN-based rating prediction. We have conducted extensive experiments over eight benchmark datasets to examine the relative superiority of the proposed KGMC method, compared to state-of-the-art baselines. Additional experiments and case studies have been also conducted to demonstrate the transferability as well as keyword-based explainability of KGMC. Our findings highlight the practical advantages of our model for recommender systems and support its effectiveness in inductive graph-based link prediction.
Donghee Han 0001, Keejun Han, Mun Yong Yi
Eng. Appl. Artif. Intell.1
2023 Context-aware Inductive Graph Matrix Completion with Sentence BERT
abstract
Existing graph neural network (GNN) based recommendation models depend highly on initial node features and graph structures. Most prior studies do not support the use of additional information on user-item interactions and adopt a transductive method, which has limitations for actual web service. To overcome these problems, we propose Context-aware Inductive Graph Matrix Completion (CGMC), which utilizes context vectors that represent user-item interactions and user-item bipartite graph structures in an inductive manner. Our method constructs context vectors from the review, timestamp, and rating and uses them as edge features of the user-item bipartite graph. Relations between homogeneous nodes are also constructed based on common neighbors. Our model combines context vectors and relational information using Context-aware Graph Attention Networks and Edge Fusion Graph Convolutional Networks. We conducted extensive experiments using six real-world datasets, and the results show that the proposed model achieves superior performances over other competing models. Furthermore, we analyzed the inductive characteristics of CGMC through the cross-domain transferability. The source code is available in the repository at [https://github.com/venzino-han/CGMC_SBERT].
Donghee Han 0001, Keejun Han, Mun Yong Yi
ICWS1
2023 Temporal enhanced inductive graph knowledge tracing
Donghee Han 0001, Keejun Han, Mun Yong Yi
Appl. Intell.1