Sunjae Kwon

dblp:207/9976 · DBLP profile ↗
← Back
13ranked-venue papers
10as first author
10since 2021 · last 2025
0000-0002-5425-6779ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 DischargeSim: A Simulation Benchmark for Educational Doctor-Patient Communication at Discharge
abstract
Discharge communication is a critical yet underexplored component of patient care, where the goal shifts from diagnosis to education.While recent large language model (LLM) benchmarks emphasize in-visit diagnostic reasoning, they fail to evaluate models' ability to support patients after the visit.We introduce DischargeSim, a novel benchmark that evaluates LLMs on their ability to act as personalized discharge educators.DischargeSim simulates post-visit, multi-turn conversations between LLM-driven DoctorAgents and Pa-tientAgents with diverse psychosocial profiles (e.g., health literacy, education, emotion).Interactions are structured across six clinically grounded discharge topics and assessed along three axes: (1) dialogue quality via automatic and LLM-as-judge evaluation, (2) personalized document generation including free-text summaries and structured AHRQ checklists, and ( 3) patient comprehension through a downstream multiple-choice exam.Experiments across 18 LLMs reveal significant gaps in discharge education capability, with performance varying widely across patient profiles.Notably, model size does not always yield better education outcomes, highlighting trade-offs in strategy use and content prioritization.DischargeSim offers a first step toward benchmarking LLMs in post-visit clinical education and promoting equitable, personalized patient support. 1 .* indicates equal contribution 1 The source code is released at: https://github.com/ michaels6060/DischargeSim with CC-BY-NC 4.0 license.Discharge notes …… -Discharge Diagnoses -Discharge Medications -Discharge Condition -Discharge Instructions -Follow-up Plan : You'll be taking furosemide 40 mg by mouth every morning-it helps reduce fluid overload and also lowers your blood pressure.But because it can make you urinate more and lower potassium, you're also on potassium chloride 10 mEq daily.I took the 40 mg furosemide today and felt lightheaded.Should I be worried?: It's not uncommon.Make sure to take the 40 mg dose after breakfast, not on an empty stomach.Also, rise slowly from sitting or lying down.If the dizziness continues or worsens, call us-we may lower the dose or adjust your schedule.…
Zonghai Yao, Michael Sun, Won Seok Jang, Sunjae Kwon, Soie Kwon, Hong Yu 0001
EMNLP4
2024 ODD: A Benchmark Dataset for the Natural Language Processing Based Opioid Related Aberrant Behavior Detection
abstract
Opioid related aberrant behaviors (ORABs) present novel risk factors for opioid overdose. This paper introduces a novel biomedical natural language processing benchmark dataset named ODD, for ORAB Detection Dataset. ODD is an expert-annotated dataset designed to identify ORABs from patients' EHR notes and classify them into nine categories; 1) Confirmed Aberrant Behavior, 2) Suggested Aberrant Behavior, 3) Opioids, 4) Indication, 5) Diagnosed opioid dependency, 6) Benzodiazepines, 7) Medication Changes, 8) Central Nervous System-related, and 9) Social Determinants of Health. We explored two state-of-the-art natural language processing models (fine-tuning and prompt-tuning approaches) to identify ORAB. Experimental results show that the prompt-tuning models outperformed the fine-tuning models in most categories and the gains were especially higher among uncommon categories (Suggested Aberrant Behavior, Confirmed Aberrant Behaviors, Diagnosed Opioid Dependence, and Medication Change). Although the best model achieved the highest 88.17% on macro average area under precision recall curve, uncommon classes still have a large room for performance improvement. ODD is publicly available.
Sunjae Kwon, Weisong Liu, Emily Druhl, Minhee L. Sung, Joel I. Reisman, Robert D. Kerns, William Becker, Hong Yu 0001
NAACL-HLT1
2023 Multi-Label Few-Shot ICD Coding as Autoregressive Generation with Prompt
abstract
Automatic International Classification of Diseases (ICD) coding aims to assign multiple ICD codes to a medical note with an average of 3,000+ tokens. This task is challenging due to the high-dimensional space of multi-label assignment (155,000+ ICD code candidates) and the long-tail challenge - Many ICD codes are infrequently assigned yet infrequent ICD codes are important clinically. This study addresses the long-tail challenge by transforming this multi-label classification task into an autoregressive generation task. Specifically, we first introduce a novel pretraining objective to generate free text diagnosis and procedure descriptions using the SOAP structure, the medical logic physicians use for note documentation. Second, instead of directly predicting the high dimensional space of ICD codes, our model generates the lower dimension of text descriptions, which then infer ICD codes. Third, we designed a novel prompt template for multi-label classification. We evaluate our Generation with Prompt (GP) model with the benchmark of all code assignment (MIMIC-III-full) and few shot ICD code assignment evaluation benchmark (MIMIC-III-few). Experiments on MIMIC-III-few show that our model performs with a marco F1 30.2, which substantially outperforms the previous MIMIC-III-full SOTA model (marco F1 4.3) and the model specifically designed for few/zero shot setting (marco F1 18.7). Finally, we design a novel ensemble learner, a cross attention reranker with prompts, to integrate previous SOTA and our best few-shot coding predictions. Experiments on MIMIC-III-full show that our ensemble learner substantially improves both macro and micro F1, from 10.4 to 14.6 and from 58.2 to 59.1, respectively.
Zhichao Yang 0001, Sunjae Kwon, Zonghai Yao, Hong Yu 0001
AAAI2
2023 Vision Meets Definitions: Unsupervised Visual Word Sense Disambiguation Incorporating Gloss Information
abstract
Visual Word Sense Disambiguation (VWSD) is a task to find the image that most accurately depicts the correct sense of the target word for the given context.Previously, image-text matching models often suffered from recognizing polysemous words.This paper introduces an unsupervised VWSD approach that uses gloss information of an external lexical knowledge-base, especially the sense definitions.Specifically, we suggest employing Bayesian inference to incorporate the sense definitions when sense information of the answer is not provided.In addition, to ameliorate the out-of-vocabulary (OOV) issue, we propose a context-aware definition generation with GPT-3.Experimental results show that VWSD performance increased significantly with our Bayesian inference-based approach.In addition, our context-aware definition generation achieved prominent performance improvement in OOV examples exhibiting better performance than the existing definition generation method.
Sunjae Kwon, Rishabh Garodia, Minhwa Lee, Zhichao Yang 0001, Hong Yu 0001
ACL (1)1
2023 Exploring LLM-based Automated Repairing of Ansible Script in Edge-Cloud Infrastructures
abstract
Edge-Cloud system requires massive infrastructures located in closer to the user to minimize latencies in handling Big data. Ansible is one of the most popular Infrastructure as Code (IaC) tools crucial for deploying these infrastructures of the Edge-cloud system. However, Ansible also consists of code, and its code quality is critical in ensuring the delivery of high-quality services within the Edge-Cloud system. On the other hand, the Large Langue Model (LLM) has performed remarkably on various Software Engineering (SE) tasks in recent years. One such task is Automated Program Repairing (APR), where LLMs assist developers in proposing code fixes for identified bugs. Nevertheless, prior studies in LLM-based APR have predominantly concentrated on widely used programming languages (PL), such as Java and C, and there has yet to be an attempt to apply it to Ansible. Hence, we explore the applicability of LLM-based APR on Ansible. We assess LLMs’ performance (ChatGPT and Bard) on 58 Ansible script revision cases from Open Source Software (OSS). Our findings reveal promising prospects, with LLMs generating helpful responses in 70% of the sampled cases. Nonetheless, further research is necessary to harness this approach’s potential fully.
Sunjae Kwon, Sungu Lee, Taehyoun Kim, Duksan Ryu, Jongmoon Baik
J. Web Eng.1
2023 Pre-trained Model-based Software Defect Prediction for Edge-cloud Systems
abstract
Edge-cloud computing is a distributed computing infrastructure that brings computation and data storage with low latency closer to clients. As interest in edge-cloud systems grows, research on testing the systems has also been actively studied. However, as with traditional systems, the amount of resources for testing is always limited. Thus, we suggest a function-level just-in-time (JIT) software defect prediction (SDP) model based on a pre-trained model to address the limitation by prioritizing the limited testing resources for the defect-prone functions. The pre-trained model is a transformer-based deep learning model trained on a large corpus of code snippets, and the fine-tuned pre-trained model can provide the defect proneness for the changed functions at a commit level. We evaluate the performance of the three popular pre-trained models (i.e., CodeBERT, GraphCodeBERT, UniXCoder) on edge-cloud systems in within-project and cross-project environments. To the best of our knowledge, it is the first attempt to analyse the performance of the three pre-trained model-based SDP models for edge-cloud systems. As a result, we can confirm that UniXCoder showed the best performance among the three in the WPDP environment. However, we also confirm that additional research is necessary to apply the SDP models to the CPDP environment.
Sunjae Kwon, Sungu Lee, Duksan Ryu, Jongmoon Baik
J. Web Eng.1
2023 An effective approach to improve the performance of eCPDP (early cross-project defect prediction) via data-transformation and parameter optimization
Sunjae Kwon, Duksan Ryu, Jongmoon Baik
Softw. Qual. J.1
2022 MedJEx: A Medical Jargon Extraction Model with Wiki's Hyperlink Span and Contextualized Masked Language Model Score
abstract
This paper proposes a new natural language processing (NLP) application for identifying medical jargon terms potentially difficult for patients to comprehend from electronic health record (EHR) notes.We first present a novel and publicly available dataset with expertannotated medical jargon terms from 18K+ EHR note sentences (M edJ).Then, we introduce a novel medical jargon extraction (M edJEx) model which has been shown to outperform existing state-of-the-art NLP models.First, MedJEx improved the overall performance when it was trained on an auxiliary Wikipedia hyperlink span dataset, where hyperlink spans provide additional Wikipedia articles to explain the spans (or terms), and then fine-tuned on the annotated MedJ data.Secondly, we found that a contextualized masked language model score was beneficial for detecting domain-specific unfamiliar jargon terms.Moreover, our results show that training on the auxiliary Wikipedia hyperlink span datasets improved six out of eight biomedical named entity recognition benchmark datasets.MedJEx is publicly available 1 .
Sunjae Kwon, Zonghai Yao, Harmon S. Jordan, David A. Levy, Brian Corner, Hong Yu 0001
EMNLP1
2021 eCPDP: Early Cross-Project Defect Prediction
abstract
Cross-project Defect Prediction (CPDP) aims to build a defect prediction model to recognize target project's defective modules by utilizing other source project's historical data. In addition, Transfer Learning (TL) has been widely applied at CPDP to improve prediction performance by alleviating the data distribution discrepancy between the source and the target project. However, existing TL-based CPDP techniques are not applicable at the unit testing phase since they require the entire historical target project data for TL. As a result, they lose a chance of increasing the product's reliability in the unit testing phase by applying the prediction results to identify defects. Thus, the objective of this paper is to apply prediction results at the unit testing phase. To this end, we propose an early CPDP model (eCPDP) which is TL-based CPDP technique using Singular Value Decomposition applicable at the unit testing phase. We compare the performance of eCPDP with state-of-the-art TL-based CPDP techniques on effort-unaware and effort-aware performance metrics over 17 project datasets. Experimental result demonstrates that eCPDP executed during the unit testing stage is one of the best techniques compared to baselines executed after the unit testing stage on both types of metrics. Thus, we show that eCPDP is an applicable CPDP model at the unit testing phase, and it can help practitioners find and fix defects in an earlier phase than other TL-based CPDP techniques.
Sunjae Kwon, Duksan Ryu, Jongmoon Baik
QRS1
2021 Word sense disambiguation based on context selection using knowledge-based word similarity
Sunjae Kwon, Dongsuk Oh, Youngjoong Ko
Inf. Process. Manag.1
2019 Effective vector representation for the Korean named-entity recognition
Sunjae Kwon, Youngjoong Ko, Jungyun Seo
Pattern Recognit. Lett.1
2018 Word Sense Disambiguation Based on Word Similarity Calculation Using Word Vector Representation from a Knowledge-based Graph
abstract
Word sense disambiguation (WSD) is the task to determine the word sense according to its context. Many existing WSD studies have been using an external knowledge-based unsupervised approach because it has fewer word set constraints than supervised approaches requiring training data. In this paper, we propose a new WSD method to generate the context of an ambiguous word by using similarities between an ambiguous word and words in the input document. In addition, to leverage our WSD method, we further propose a new word similarity calculation method based on the semantic network structure of BabelNet. We evaluate the proposed methods on the SemEval-13 and SemEval-15 for English WSD dataset. Experimental results demonstrate that the proposed WSD method significantly improves the baseline WSD method. Furthermore, our WSD system outperforms the state-of-the-art WSD systems in the Semeval-13 dataset. Finally, it has higher performance than the state-of-the-art unsupervised knowledge-based WSD system in the average performance of both datasets.
Dongsuk Oh, Sunjae Kwon, Kyungsun Kim, Youngjoong Ko
COLING2
2017 A Robust Named-Entity Recognition System Using Syllable Bigram Embedding with Eojeol Prefix Information
abstract
Korean named-entity recognition (NER) systems have been developed mainly on the morphological-level, and they are commonly based on a pipeline framework that identifies named-entities (NEs) following the morphological analysis. However, this framework can mean that the performance of NER systems is degraded, because errors from the morphological analysis propagate into NER systems. This paper proposes a novel syllable-level NER system, which does not require a morphological analysis and can achieve a similar or better performance compared with the morphological-level NER systems. In addition, because the proposed system does not require a morphological analysis step, its processing speed is about 1.9 times faster than those of the previous morphological-level NER systems.
Sunjae Kwon, Youngjoong Ko, Jungyun Seo
CIKM1