Misuk Kim

dblp:61/6756 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0002-8623-3088ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 3 first-author · 13 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 BIND: A Bidirectionally Aligned Next-token Denoising Framework for Fast and Lightweight Deobfuscation of Harmful Web Text
abstract
Harmful online content, including hate speech, fraud, and phishing, is increasingly disseminated in obfuscated forms designed to evade detection. This creates an urgent need for accurate and efficient real-time de-obfuscation methods to protect users and maintain trust. Existing obfuscation detection methods rely on large auto-regressive models and byte-level fallback tokenizers, which are hindered by slow inference speeds and face difficulties in handling graphemes with multiple code points and out-of-vocabulary (OOV) processing. This study proposes Bidirectionally Aligned Next-Token Denoising ( BIND ), which integrates character-level token alignment with a novel attention technique to enable precise and efficient corrections at fixed positions. Experiments conducted on a public dataset of obfuscated harmful text demonstrate that BIND outperforms existing methods. BIND has shown strong robustness against various text-based visual, phonetic, and semantic perturbations, proving particularly resilient against emojis and other OOV elements. This research highlights how a task-specific small language model can outperform larger ones, offering a practical solution for real-time harmful content mitigation and contributing to the development of a safer and more responsible web.
Misuk Kim
WWW2
2026 BearGen: LLM-guided signal generation framework for bearing fault diagnosis
abstract
Signal data are essential for condition monitoring, fault diagnosis, and decision-making across industrial domains, and research leveraging signal data has been actively pursued in areas such as healthcare and manufacturing. However, acquiring such data is costly and difficult due to factors such as the risk of equipment damage, the need for expert labeling, and the scarcity of fault data. Moreover, collected data often contain sensitive operational information, making sharing difficult, and enterprises are restricted from using high-performance models hosted on external servers due to security concerns. To address these challenges, we propose BearGen , a novel framework that combines the strong generative capabilities of Large Language Models (LLMs) with the precise data distribution learning of diffusion models to synthesize high-quality signal data in on-premise environments. BearGen first employs an LLM to generate descriptions of existing signals and then conditions a description-guided diffusion model on these descriptions to generate high-quality synthetic signals. We evaluated BearGen on eight publicly available bearing fault diagnosis datasets, and the results showed superior performance compared to existing approaches. In addition, we experimentally validated the reliability and usefulness of the generated signal descriptions. Further experiments under conditions simulating real industrial environments — such as limited data availability and severe data imbalance — verified the practical applicability of the framework. By operating in on-premise environments, BearGen resolves data security concerns while alleviating data scarcity and imbalance. Furthermore, by providing natural language descriptions, it enhances interpretability and offers significant potential for decision support in real-world industrial applications.
Hyuna Jeon, Uiin Kim, Misuk Kim
Adv. Eng. Informatics4
2026 A scalable unsupervised framework for multi-aspect labeling of multilingual and multi-domain review data
Jiin Park, Misuk Kim
Knowl. Based Syst.2
2026 RTKD: Responsive Teacher Knowledge Distillation with heterogeneous students
Geonyeong Son, Misuk Kim
Knowl. Based Syst.2
2025 Multi-level prompting: Enhancing model performance through hierarchical instruction integration
Geonyeong Son, Misuk Kim
Knowl. Based Syst.2
2024 Hierarchical Graph Convolutional Network Approach for Detecting Low-Quality Documents
abstract
Consistency within a document is a crucial feature indicative of its quality. Recently, within the vast amount of information produced across various media, there exists a significant number of low-quality documents that either lack internal consistency or contain content utterly unrelated to their headlines. Such low-quality documents induce fatigue in readers and undermine the credibility of the media source that provided them. Consequently, research to automatically detect these low-quality documents based on natural language processing is imperative. In this study, we introduce a hierarchical graph convolutional network (HGCN) that can detect internal inconsistencies within a document and incongruences between the title and body. Moreover, we constructed the Inconsistency Dataset, leveraging published news data and its meta-data, to train our model to detect document inconsistencies. Experimental results demonstrated that the HGCN achieved superior performance with an accuracy of 91.20% on our constructed Inconsistency Dataset, outperforming other comparative models. Additionally, on the publicly available incongruent-related dataset, the proposed methodology demonstrated a performance of 92.00%, validating its general applicability. Finally, an ablation study further confirmed the significant impact of meta-data utilization on performance enhancement. We anticipate that our model can be universally applied to detect and filter low-quality documents in the real world.
Joonwon Jang, Misuk Kim
LREC/COLING3
2024 A simple and efficient dialogue generation model incorporating commonsense knowledge
Geonyeong Son, Misuk Kim
Expert Syst. Appl.2
2024 DSTEA: Improving Dialogue State Tracking via Entity Adaptive pre-training
Yukyung Lee, Takyoung Kim, Hoonsang Yoon, Pilsung Kang 0001, Junseong Bang, Misuk Kim
Knowl. Based Syst.6
2023 ESG information extraction with cross-sectoral and multi-source adaptation based on domain-tuned language models
Misuk Kim
Expert Syst. Appl.2
2023 Dialogue Specific Pre-training Tasks for Improved Dialogue State Tracking
Jinwon An, Misuk Kim
Neural Process. Lett.2
2022 Detecting incongruent news headlines with auxiliary textual information
Joonwon Jang, Yoon-Sik Cho, Misuk Kim
Expert Syst. Appl.4
2022 Accurate and prompt answering framework based on customer reviews and question-answer pairs
Eun Kim, Hyejung Yoon, Misuk Kim
Expert Syst. Appl.4
2022 Domain-Slot Relationship Modeling Using a Pre-Trained Language Encoder for Multi-Domain Dialogue State Tracking
abstract
Dialogue state tracking for multi-domain dialogues is challenging because the model should be able to track dialogue states across multiple domains and slots. As using pre-trained language models is the de facto standard for natural language processing tasks, many recent studies use them to encode the dialogue context for predicting the dialogue states. Model architectures that have certain inductive biases for modeling the relationship among different domain-slot pairs are also emerging. Our work is based on these research approaches on multi-domain dialogue state tracking. We propose a model architecture that effectively models the relationship among domain-slot pairs using a pre-trained language encoder. Inspired by the way the special [CLS] token in BERT is used to aggregate the information of the whole sequence, we use multiple special tokens for each domain-slot pair that encodes information corresponding to its domain and slot. The special tokens are run together with the dialogue context through the pre-trained language encoder, which effectively models the relationship among different domain-slot pairs. Our experimental results on the datasets MultiWOZ-2.0 and MultiWOZ-2.1 show that our model outperforms other models with the same setting. Our ablation studies incorporate three main parts. The first component shows the effectiveness of our approach exploiting the relationship modeling. The second component compares the effect of using different pre-trained language encoders. The final component involves comparing different initialization methods that could be used for the special tokens. Qualitative analysis of the attention map of the pre-trained language encoder shows that our special tokens encode relevant information through the encoding process by attending to each other.
Jinwon An, Sungzoon Cho, Junseong Bang, Misuk Kim
IEEE ACM Trans. Audio Speech Lang. Process.4
2021 A data mining framework for financial prediction
Misuk Kim
Expert Syst. Appl.1
2021 Adaptive trading system integrating machine learning and back-testing: Korean bond market case
Misuk Kim
Expert Syst. Appl.1
2018 Stock price prediction through sentiment analysis of corporate disclosures using distributed representation
abstract
Many researches have exploited textual data, such as news, online blogs, and financial reports, in order to predict stock price movements effectively. Previous studies formed the task as a classification problem predicting upward or downward movement of stock prices from text documents. Such an app roach, however, may be deemed inappropriate when combined with sentiment analysis. In financial documents, same words may convey different sentiments across different sectors; if documents from multiple sectors are learned simultaneously, performance can deteriorate. Therefore, we conducted sentiment analysis of 8-K financial reports of firms sector by sector. In particular, we also employed distributed representation for predicting stock price movements. Experiment results show that our approach improves prediction performance by 25.4% over the baseline model, and that the direction of post-announcement stock price movements shifts accordingly with the polarity of the sentiment of reports. Not only does our model improve predictability, but also provides visualizations, which may assist agents actively trading in the field with understanding the drivers for the observed stock movements. The two main aspects of our model, predictability and interpretability, will provide meaningful insights to help decision-makers in the industry with time-split trading decisions or data-driven detection of promising companies.
Misuk Kim, Eunjeong L. Park, Sungzoon Cho
Intell. Data Anal.1
2005 The Effects of Interactive Video Printer Technology on Encoding Strategies in Young Children's Episodic Memory
Misuk Kim
ACII1