Hyunjae Kim

dblp:138/1746 · DBLP profile ↗
← Back
20ranked-venue papers
3as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 MED-COREASONER: Reducing Language Disparities in Medical Reasoning via Language-Informed Co-Reasoning
abstract
While reasoning-enhanced large language models perform strongly on English medical tasks, a persistent multilingual gap remains, with substantially weaker reasoning in local languages, limiting equitable global medical deployment. To bridge this gap, we introduce Med-CoReasoner, a language-informed co-reasoning framework that elicits parallel English and local-language reasoning, abstracts them into structured concepts, and integrates local clinical knowledge into an English logical scaffold via concept-level alignment and retrieval. This design combines the structural robustness of English reasoning with the practice-grounded expertise encoded in local languages. To evaluate multilingual medical reasoning beyond multiple-choice settings, we construct MultiMed-X, a benchmark covering seven languages with expert-annotated long-form question answering and natural language inference tasks, comprising 350 instances per language. Experiments across three benchmarks show that Med-CoReasoner improves multilingual reasoning performance by an average of 5%, with particularly substantial gains in low-resource languages. Moreover, model distillation and expert evaluation analysis further confirm that Med-CoReasoner produces clinically sound and culturally grounded reasoning traces.
Sherry T. Tong, Jiwoong Sohn, Ding Xia, Piyalitt Ittichaiwong, Kanyakorn Veerakanjana, Hyunjae Kim, Qingyu Chen 0001, Edison Marrese-Taylor, Kazuma Kobayashi, Akiko Aizawa, Irene Li
ACL (1)9
2025 Med-PRM: Medical Reasoning Models with Stepwise, Guideline-verified Process Rewards
abstract
Jaehoon Yun, Jiwoong Sohn, Jungwoo Park, Hyunjae Kim, Xiangru Tang, Daniel Shao, Yong Hoe Koo, Ko Minhyeok, Qingyu Chen, Mark Gerstein, Michael Moor, Jaewoo Kang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Jaehoon Yun, Jiwoong Sohn, Jungwoo Park, Hyunjae Kim, Xiangru Tang, Daniel Shao, Yonghoe Koo, Minhyeok Ko, Qingyu Chen 0001, Mark Gerstein, Michael Moor, Jaewoo Kang
EMNLP4
2025 ETHIC: Evaluating Large Language Models on Long-Context Tasks with High Information Coverage
abstract
Taewhoo Lee, Chanwoong Yoon, Kyochul Jang, Donghyeon Lee, Minju Song, Hyunjae Kim, Jaewoo Kang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Taewhoo Lee, Chanwoong Yoon, Kyochul Jang, Minju Song, Hyunjae Kim, Jaewoo Kang
NAACL (Long Papers)6
2025 Rationale-Guided Retrieval Augmented Generation for Medical Question Answering
abstract
Jiwoong Sohn, Yein Park, Chanwoong Yoon, Sihyeon Park, Hyeon Hwang, Mujeen Sung, Hyunjae Kim, Jaewoo Kang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Jiwoong Sohn, Yein Park, Chanwoong Yoon, Sihyeon Park, Hyeon Hwang, Mujeen Sung, Hyunjae Kim, Jaewoo Kang
NAACL (Long Papers)7
2024 CookingSense: A Culinary Knowledgebase with Multidisciplinary Assertions
abstract
This paper introduces CookingSense, a descriptive collection of knowledge assertions in the culinary domain extracted from various sources, including web data, scientific papers, and recipes, from which knowledge covering a broad range of aspects is acquired. CookingSense is constructed through a series of dictionary-based filtering and language model-based semantic filtering techniques, which results in a rich knowledgebase of multidisciplinary food-related assertions. Additionally, we present FoodBench, a novel benchmark to evaluate culinary decision support systems. From evaluations with FoodBench, we empirically prove that CookingSense improves the performance of retrieval augmented language models. We also validate the quality and variety of assertions in CookingSense through qualitative analysis.
Donghee Choi, Keonwoo Kim 0002, Donghyeon Park, Mujeen Sung, Hyunjae Kim, Jaewoo Kang
LREC/COLING5
2024 Leveraging Adapter for Parameter-Efficient ASR Encoder
Kyuhong Shim, Jinkyu Lee 0004, Hyunjae Kim
INTERSPEECH3
2024 Augmenting biomedical named entity recognition with general-domain resources
abstract
OBJECTIVE: Training a neural network-based biomedical named entity recognition (BioNER) model usually requires extensive and costly human annotations. While several studies have employed multi-task learning with multiple BioNER datasets to reduce human effort, this approach does not consistently yield performance improvements and may introduce label ambiguity in different biomedical corpora. We aim to tackle those challenges through transfer learning from easily accessible resources with fewer concept overlaps with biomedical datasets. METHODS: We proposed GERBERA, a simple-yet-effective method that utilized general-domain NER datasets for training. We performed multi-task learning to train a pre-trained biomedical language model with both the target BioNER dataset and the general-domain dataset. Subsequently, we fine-tuned the models specifically for the BioNER dataset. RESULTS: We systematically evaluated GERBERA on five datasets of eight entity types, collectively consisting of 81,410 instances. Despite using fewer biomedical resources, our models demonstrated superior performance compared to baseline models trained with additional BioNER datasets. Specifically, our models consistently outperformed the baseline models in six out of eight entity types, achieving an average improvement of 0.9% over the best baseline performance across eight entities. Our method was especially effective in amplifying performance on BioNER datasets characterized by limited data, with a 4.7% improvement in F1 scores on the JNLPBA-RNA dataset. CONCLUSION: This study introduces a new training method that leverages cost-effective general-domain NER datasets to augment BioNER models. This approach significantly improves BioNER model performance, making it a valuable asset for scenarios with scarce or costly biomedical datasets. We make data, codes, and models publicly available via https://github.com/qingyu-qc/bioner_gerbera.
Hyunjae Kim, Chih-Hsuan Wei, Jaewoo Kang, Zhiyong Lu, Hua Xu 0001, Qingyu Chen 0001
J. Biomed. Informatics2
2023 LIQUID: A Framework for List Question Answering Dataset Generation
abstract
Question answering (QA) models often rely on large-scale training datasets, which necessitates the development of a data generation framework to reduce the cost of manual annotations. Although several recent studies have aimed to generate synthetic questions with single-span answers, no study has been conducted on the creation of list questions with multiple, non-contiguous spans as answers. To address this gap, we propose LIQUID, an automated framework for generating list QA datasets from unlabeled corpora. We first convert a passage from Wikipedia or PubMed into a summary and extract named entities from the summarized text as candidate answers. This allows us to select answers that are semantically correlated in context and is, therefore, suitable for constructing list questions. We then create questions using an off-the-shelf question generator with the extracted entities and original passage. Finally, iterative filtering and answer expansion are performed to ensure the accuracy and completeness of the answers. Using our synthetic data, we significantly improve the performance of the previous best list QA models by exact-match F1 scores of 5.0 on MultiSpanQA, 1.9 on Quoref, and 2.8 averaged across three BioASQ benchmarks.
Seongyun Lee, Hyunjae Kim, Jaewoo Kang
AAAI2
2023 Automatic Creation of Named Entity Recognition Datasets by Querying Phrase Representations
abstract
Most weakly supervised named entity recognition (NER) models rely on domain-specific dictionaries provided by experts.This approach is infeasible in many domains where dictionaries do not exist.While a phrase retrieval model was used to construct pseudo-dictionaries with entities retrieved from Wikipedia automatically in a recent study, these dictionaries often have limited coverage because the retriever is likely to retrieve popular entities rather than rare ones.In this study, we present a novel framework, HighGEN, that generates NER datasets with high-coverage pseudo-dictionaries.Specifically, we create entity-rich dictionaries with a novel search method, called phrase embedding search, which encourages the retriever to search a space densely populated with various entities.In addition, we use a new verification process based on the embedding distance between candidate entity mentions and entity types to reduce the false-positive noise in weak labels generated by high-coverage dictionaries.We demonstrate that HighGEN outperforms the previous best model by an average F1 score of 4.7 across five NER benchmark datasets.
Hyunjae Kim, Jaehyo Yoo, Seunghyun Yoon 0002, Jaewoo Kang
ACL (1)1
2022 Simple Questions Generate Named Entity Recognition Datasets
abstract
Recent named entity recognition (NER) models often rely on human-annotated datasets, requiring the significant engagement of professional knowledge on the target domain and entities.This research introduces an ask-to-generate approach that automatically generates NER datasets by asking questions in simple natural language to an open-domain question answering system (e.g., "Which disease?").Despite using fewer in-domain resources, our models, solely trained on the generated datasets, largely outperform strong low-resource models by an average F1 score of 19.4 for six popular NER benchmarks.Furthermore, our models provide competitive performance with rich-resource models that additionally leverage in-domain dictionaries provided by domain experts.In few-shot NER, we outperform the previous best model by an F1 score of 5.2 on three benchmarks and achieve new state-of-the-art performance.The code and datasets are available at https://github.com/dmis-lab/GeNER.
Hyunjae Kim, Jaehyo Yoo, Seunghyun Yoon 0002, Jinhyuk Lee, Jaewoo Kang
EMNLP1
2021 Learn to Resolve Conversational Dependency: A Consistency Training Framework for Conversational Question Answering
abstract
Gangwoo Kim, Hyunjae Kim, Jungsoo Park, Jaewoo Kang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Gangwoo Kim, Hyunjae Kim, Jungsoo Park, Jaewoo Kang
ACL/IJCNLP (1)2
2021 "Killing Me" Is Not a Spoiler: Spoiler Detection Model using Graph Neural Networks with Dependency Relation-Aware Attention Mechanism
abstract
Several machine learning-based spoiler detection models have been proposed recently to protect users from spoilers on review websites.Although dependency relations between context words are important for detecting spoilers, current attention-based spoiler detection models are insufficient for utilizing dependency relations.To address this problem, we propose a new spoiler detection model called SDGNN that is based on syntax-aware graph neural networks.In the experiments on two realworld benchmark datasets, we show that our SDGNN outperforms the existing spoiler detection models.
Buru Chang, Inggeol Lee, Hyunjae Kim, Jaewoo Kang
EACL3
2021 Exploring the spatial reasoning ability of neural models in human IQ tests
Hyunjae Kim, Yookyung Koh, Jinheon Baek, Jaewoo Kang
Neural Networks1
2020 Look at the First Sentence: Position Bias in Question Answering
abstract
Many extractive question answering models are trained to predict start and end positions of answers.The choice of predicting answers as positions is mainly due to its simplicity and effectiveness.In this study, we hypothesize that when the distribution of the answer positions is highly skewed in the training set (e.g., answers lie only in the k-th sentence of each passage), QA models predicting answers as positions can learn spurious positional cues and fail to give answers in different positions.We first illustrate this position bias in popular extractive QA models such as BiDAF and BERT and thoroughly examine how position bias propagates through each layer of BERT.To safely deliver position information without position bias, we train models with various de-biasing methods including entropy regularization and bias ensembling.Among them, we found that using the prior distribution of answer positions as a bias model is very effective at reducing position bias, recovering the performance of BERT from 37.48% to 81.64% when trained on a biased SQuAD dataset.
Miyoung Ko, Jinhyuk Lee, Hyunjae Kim, Gangwoo Kim, Jaewoo Kang
EMNLP (1)3
2019 Predicting Multiple Demographic Attributes with Task Specific Embedding Transformation and Attention Network
abstract
Most companies utilize demographic information to develop their strategy in a market. However, such information is not available to most retail companies. Several studies have been conducted to predict the demographic attributes of users from their transaction histories, but they have some limitations. First, they focused on parameter sharing to predict all attributes but capturing task-specific features is also important in multi-task learning. Second, they assumed that all transactions are equally important in predicting demographic attributes. However, some transactions are more useful than others for predicting a certain attribute. Furthermore, decision making process of models cannot be interpreted as they work in a black-box manner. To address the limitations, we propose an Embedding Transformation Network with Attention (ETNA) model which shares representations at the bottom of the model structure and transforms them to task-specific representations using a simple linear transformation method. In addition, we can obtain more informative transactions for predicting certain attributes using the attention mechanism. The experimental results show that our model outperforms the previous models on all tasks. In our qualitative analysis, we show the visualization of attention weights, which provides business managers with some useful insights.
Raehyun Kim, Hyunjae Kim, Janghyuk Lee, Jaewoo Kang
SDM2
2018 Ranking Paragraphs for Improving Answer Recall in Open-Domain Question Answering
abstract
Recently, open-domain question answering (QA) has been combined with machine comprehension models to find answers in a large knowledge source.As open-domain QA requires retrieving relevant documents from text corpora to answer questions, its performance largely depends on the performance of document retrievers.However, since traditional information retrieval systems are not effective in obtaining documents with a high probability of containing answers, they lower the performance of QA systems.Simply extracting more documents increases the number of irrelevant documents, which also degrades the performance of QA systems.In this paper, we introduce Paragraph Ranker which ranks paragraphs of retrieved documents for a higher answer recall with less noise.We show that ranking paragraphs and aggregating answers using Paragraph Ranker improves performance of open-domain QA pipeline on the four opendomain QA datasets by 7.8% on average.
Jinhyuk Lee, Seongjun Yun, Hyunjae Kim, Miyoung Ko, Jaewoo Kang
EMNLP3
2018 Liver Lesion Detection from Weakly-Labeled Multi-phase CT Volumes with a Grouped Single Shot MultiBox Detector
Sang-gil Lee, Jae Seok Bae, Hyunjae Kim, Sungroh Yoon
MICCAI (2)3
2018 A Deep Neural Spoiler Detection Model Using a Genre-Aware Attention Mechanism
Buru Chang, Hyunjae Kim, Raehyun Kim, Deahan Kim, Jaewoo Kang
PAKDD (1)2
2017 Transfer Learning for Deep Learning on Graph-Structured Data
abstract
Graphs provide a powerful means for representing complex interactions between entities. Recently, new deep learning approaches have emerged for representing and modeling graph-structured data while the conventional deep learning methods, such as convolutional neural networks and recurrent neural networks, have mainly focused on the grid-structured inputs of image and audio. Leveraged by representation learning capabilities, deep learning-based techniques can detect structural characteristics of graphs, giving promising results for graph applications. In this paper, we attempt to advance deep learning for graph-structured data by incorporating another component: transfer learning. By transferring the intrinsic geometric information learned in the source domain, our approach can construct a model for a new but related task in the target domain without collecting new data and without training a new model from scratch. We thoroughly tested our approach with large-scale real-world text data and confirmed the effectiveness of the proposed transfer learning framework for deep learning on graphs. According to our experiments, transfer learning is most effective when the source and target domains bear a high level of structural similarity in their graph representations.
Jaekoo Lee, Hyunjae Kim, Jongsun Lee, Sungroh Yoon
AAAI2
2017 Name Nationality Classification with Recurrent Neural Networks
abstract
Personal names tend to have many variations differing from country to country. Though there exists a large amount of personal names on the Web, nationality prediction solely based on names has not been fully studied due to its difficulties in extracting subtle character level features. We propose a recurrent neural network based model which predicts nationalities of each name using automatic feature extraction. Evaluation of Olympic record data shows that our model achieves greater accuracy than previous feature based approaches in nationality prediction tasks. We also evaluate our proposed model and baseline models on name ethnicity classification task, again achieving better or comparable performances. We further investigate the effectiveness of character embeddings used in our proposed model.
Jinhyuk Lee, Hyunjae Kim, Miyoung Ko, Donghee Choi, Jaewoo Kang
IJCAI2