JinYeong Bak

dblp:22/11519 · also Jinyeong Bak · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
13since 2021 · last 2025
0000-0002-3212-5241ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 4 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
abstract
The application scope of Large Language Models (LLMs) continues to expand, leading to increasing interest in personalized LLMs that align with human values.However, aligning these models with individual values raises significant safety concerns, as certain values may correlate with harmful information.In this paper, we identify specific safety risks associated with value-aligned LLMs and investigate the psychological principles behind these challenges.Our findings reveal two key insights.(1) Value-aligned LLMs are more prone to harmful behavior compared to non-fine-tuned models and exhibit slightly higher risks in traditional safety evaluations than other fine-tuned models.(2) These safety issues arise because value-aligned LLMs genuinely generate text according to the aligned values, which can amplify harmful outcomes.Using a dataset with detailed safety categories, we find significant correlations between value alignment and safety risks, supported by psychological hypotheses.This study offers insights into the "black box" of value alignment and proposes in-context alignment methods to enhance the safety of value-aligned LLMs. 1
Sooyung Choi, Jaehyeok Lee, Xiaoyuan Yi, Jing Yao 0003, Xing Xie 0001, JinYeong Bak
ACL (1)6
2025 Self-Training Meets Consistency: Improving LLMs' Reasoning with Consistency-Driven Rationale Evaluation
abstract
Jaehyeok Lee, Keisuke Sakaguchi, JinYeong Bak. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Jaehyeok Lee, Keisuke Sakaguchi, JinYeong Bak
NAACL (Long Papers)3
2025 Dialogue response coherency evaluation with feature sensitive negative sample using multi list-wise ranking loss
Yeongjun Hwang, Dongjun Kang, JinYeong Bak
Eng. Appl. Artif. Intell.3
2025 An intelligent marketing platform with influencer classification in social networking services
Xiaohong Yu, Jinyong Kim, Yoseop Ahn, Mose Gu, Jaehoon Jeong 0001, JinYeong Bak, Jaemin Jo
Knowl. Based Syst.6
2024 Perturb-and-Compare Approach for Detecting Out-of-Distribution Samples in Constrained Access Environments
abstract
Accessing machine learning models through remote APIs has been gaining prevalence following the recent trend of scaling up model parameters for increased performance. Even though these models exhibit remarkable ability, detecting out-of-distribution (OOD) samples remains a crucial safety concern for end users as these samples may induce unreliable outputs from the model. In this work, we propose an OOD detection framework, MixDiff, that is applicable even when the model’s parameters or its activations are not accessible to the end user. To bypass the access restriction, MixDiff applies an identical input-level perturbation to a given target sample and a similar in-distribution (ID) sample, then compares the relative difference in the model outputs of these two samples. MixDiff is model-agnostic and compatible with existing output-based OOD detection methods. We provide theoretical analysis to illustrate MixDiff’s effectiveness in discerning OOD samples that induce overconfident outputs from the model and empirically demonstrate that MixDiff consistently enhances the OOD detection performance on various datasets in vision and text domains.
Hoyoon Byun, Changdae Oh, JinYeong Bak, Kyungwoo Song
ECAI4
2024 Memoria: Resolving Fateful Forgetting Problem through Human-Inspired Memory Architecture
abstract
Making neural networks remember over the long term has been a longstanding issue. Although several external memory techniques have been introduced, most focus on retaining recent information in the short term. Regardless of its importance, information tends to be fatefully forgotten over time. We present Memoria, a memory system for artificial neural networks, drawing inspiration from humans and applying various neuroscientific and psychological theories. The experimental results prove the effectiveness of Memoria in the diverse tasks of sorting, language modeling, and classification, surpassing conventional techniques. Engram analysis reveals that Memoria exhibits the primacy, recency, and temporal contiguity effects which are characteristics of human memory.
JinYeong Bak
ICML2
2024 PEMA: An Offsite-Tunable Plug-in External Memory Adaptation for Language Models
abstract
HyunJin Kim, Young Jin Kim, JinYeong Bak. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
JinYeong Bak
NAACL-HLT3
2023 A Transformer-based Function Symbol Name Inference Model from an Assembly Language for Binary Reversing
abstract
Reverse engineering of a stripped binary has a wide range of applications, yet it is challenging mainly due to the lack of contextually useful information within. Once debugging symbols (e.g., variable names, types, function names) are discarded, recovering such information is not technically viable with traditional approaches like static or dynamic binary analysis. We focus on a function symbol name recovery, which allows a reverse engineer to gain a quick overview of an unseen binary. The key insight is that a well-developed program labels a meaningful function name that describes its underlying semantics well. In this paper, we present AsmDepictor, the Transformer-based framework that generates a function symbol name from a set of assembly codes (i.e., machine instructions), which consists of three major components: binary code refinement, model training, and inference. To this end, we conduct systematic experiments on the effectiveness of code refinement that can enhance an overall performance. We introduce the per-layer positional embedding and Unique-softmax for AsmDepictor so that both can aid to capture a better relationship between tokens. Lastly, we devise a novel evaluation metric tailored for a short description length, the Jaccard* score. Our empirical evaluation shows that the performance of AsmDepictor by far surpasses that of the state-of-the-art models up to around 400%. The best AsmDepictor model achieves an F1 of 71.5 and Jaccard* of 75.4.
JinYeong Bak, Kyunghyun Cho, Hyungjoon Koo
AsiaCCS2
2023 Conversational Emotion-Cause Pair Extraction with Guided Mixture of Experts
abstract
Emotion-Cause Pair Extraction (ECPE) task aims to pair all emotions and corresponding causes in documents.ECPE is an important task for developing human-like responses.However, previous ECPE research is conducted based on news articles, which has different characteristics compared to dialogues.To address this issue, we propose a Pair-Relationship Guided Mixture-of-Experts (PRG-MoE) model, which considers dialogue features (e.g., speaker information).PRG-MoE automatically learns relationship between utterances and advises a gating network to incorporate dialogue features in the evaluation, yielding substantial performance improvement.We employ a new ECPE dataset, which is an English dialogue dataset, with more emotion-cause pairs in documents than news articles.We also propose Cause Type Classification that classifies emotion-cause pairs according to the types of the cause of a detected emotion.For reproducing the results, we make available all our code and data 1 .
Dongjin Jeong, JinYeong Bak
EACL2
2023 From Values to Opinions: Predicting Human Behaviors and Stances Using Value-Injected Large Language Models
abstract
Being able to predict people's opinions on issues and behaviors in realistic scenarios can be helpful in various domains, such as politics and marketing.However, conducting largescale surveys like the European Social Survey to solicit people's opinions on individual issues can incur prohibitive costs.Leveraging prior research showing influence of core human values on individual decisions and actions, we propose to use value-injected large language models (LLM) to predict opinions and behaviors.To this end, we present Value Injection Method (VIM), a collection of two methodsargument generation and question answeringdesigned to inject targeted value distributions into LLMs via fine-tuning.We then conduct a series of experiments on four tasks to test the effectiveness of VIM and the possibility of using value-injected LLMs to predict opinions and behaviors of people.We find that LLMs valueinjected with variations of VIM substantially outperform the baselines.Also, the results suggest that opinions and behaviors can be better predicted using value-injected LLMs than the baseline approaches.
Dongjun Kang, Joonsuk Park, Yohan Jo, JinYeong Bak
EMNLP4
2023 It Ain't Over: A Multi-aspect Diverse Math Word Problem Dataset
abstract
The math word problem (MWP) is a complex task that requires natural language understanding and logical reasoning to extract key knowledge from natural language narratives.Previous studies have provided various MWP datasets but lack diversity in problem types, lexical usage patterns, languages, and annotations for intermediate solutions.To address these limitations, we introduce a new MWP dataset, named DMath (Diverse Math Word Problems), offering a wide range of diversity in problem types, lexical usage patterns, languages, and intermediate solutions.The problems are available in English and Korean and include an expression tree and Python code as intermediate solutions.Through extensive experiments, we demonstrate that the DMath dataset provides a new opportunity to evaluate the capability of large language models, i.e., GPT-4 only achieves about 75% accuracy on the DMath 1 dataset.
Ilwoong Baek, JinYeong Bak, Jongwuk Lee
EMNLP4
2023 Diversity Enhanced Narrative Question Generation for Storybooks
abstract
Question generation (QG) from a given context can enhance comprehension, engagement, assessment, and overall efficacy in learning or conversational environments.Despite recent advancements in QG, the challenge of enhancing or measuring the diversity of generated questions often remains unaddressed.In this paper, we introduce a multi-question generation model (mQG), which is capable of generating multiple, diverse, and answerable questions by focusing on context and questions.To validate the answerability of the generated questions, we employ a SQuAD2.0fine-tuned question answering model, classifying the questions as answerable or not.We train and evaluate mQG on the FairytaleQA dataset, a well-structured QA dataset based on storybooks, with narrative questions.We further apply a zero-shot adaptation on the TellMeWhy and SQuAD1.1 datasets.mQG shows promising results across various evaluation metrics, among strong baselines.1
Hokeun Yoon, JinYeong Bak
EMNLP2
2022 Genre-Controllable Story Generation via Supervised Contrastive Learning
abstract
While controllable text generation has received attention due to the recent advances in large-scale pre-trained language models, there is a lack of research that focuses on story-specific controllability. To address this, we present Story Control via Supervised Contrastive learning model (SCSC), to create a story conditioned on genre. For this, we design a supervised contrastive objective combined with log-likelihood objective, to capture the intrinsic differences among the stories in different genres. The results of our automated evaluation and user study demonstrate that the proposed method is effective in genre-controlled story generation.
Jin-Uk Cho, Min-Su Jeong, JinYeong Bak, Yun-Gyung Cheong
WWW3
2020 Speaker Sensitive Response Evaluation Model
abstract
Automatic evaluation of open-domain dialogue response generation is very challenging because there are many appropriate responses for a given context.Existing evaluation models merely compare the generated response with the ground truth response and rate many of the appropriate responses as inappropriate if they deviate from the ground truth.One approach to resolve this problem is to consider the similarity of the generated response with the conversational context.In this paper, we propose an automatic evaluation model based on that idea and learn the model parameters from an unlabeled conversation corpus.Our approach considers the speakers in defining the different levels of similar context.We use a Twitter conversation corpus that contains many speakers and conversations to test our evaluation model.Experiments show that our model outperforms the other existing evaluation metrics in terms of high correlation with human annotation scores.We also show that our model trained on Twitter can be applied to movie dialogues without any additional training.We provide our code and the learned parameters so that they can be used for automatic evaluation of dialogue response generation models.
JinYeong Bak, Alice Oh
ACL1
2019 Variational Hierarchical User-based Conversation Model
abstract
JinYeong Bak, Alice Oh. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
JinYeong Bak, Alice Oh
EMNLP/IJCNLP (1)1
2018 Conversational Decision Making Model for Predicting King's Decision in the Annals of the Joseon Dynasty
abstract
Styles of leaders when they make decisions in groups vary, and the different styles affect the performance of the group.To understand the key words and speakers associated with decisions, we initially formalize the problem as one of predicting leaders' decisions from discussion with group members.As a dataset, we introduce conversational meeting records from a historical corpus, and develop a hierarchical RNN structure with attention and pre-trained speaker embedding in the form of a, Conversational Decision Making Model (CDMM).The CDMM outperforms other baselines to predict leaders' final decisions from the data.We explain why CDMM works better than other methods by showing the key words and speakers discovered from the attentions as evidence.
JinYeong Bak, Alice Oh
EMNLP1
2017 Itchtector: A Wearable-based Mobile System for Managing Itching Conditions
abstract
Severe itching conditions such as eczema or atopic dermatitis can have a significant impact on one's quality of life. Unfortunately, many of these conditions cannot be cured, and the focus is often on properly controlling or managing the condition. Thus, it is important to understand or objectively monitor how one's scratching behavior changes, based on medication or treatment or environmental conditions. In this work, we explore how wearable devices can support people with itching conditions to better manage their conditions. We carried out a three-phase study with 40 participants and 2 dermatologists to understand the implications of various system features and designs. Based on interviews with patients and doctors, we incorporated medical guidelines for treatment and patients' needs in the proposed Itchtector - a smartwatch-based mobile system to monitor itching behaviors and provide objective information about the user's scratching behaviors. Using the Itchtector prototype, we evaluated performance and possible acceptance with subjects.
Jongin Lee, Dae-ki Cho, Junhong Kim, Eunji Im, JinYeong Bak, Kyung ho Lee, KwanHong Lee, John Kim 0001
CHI5
2017 Rotated Word Vector Representations and their Interpretability
abstract
Vector representation of words improves performance in various NLP tasks, but the high-dimensional word vectors are very difficult to interpret.We apply several rotation algorithms to the vector representation of words to improve the interpretability.Unlike previous approaches that induce sparsity, the rotated vectors are interpretable while preserving the expressive performance of the original vectors.Furthermore, any pre-built word vector representation can be rotated for improved interpretability.We apply rotation to skipgrams and glove and compare the expressive power and interpretability with the original vectors and the sparse overcomplete vectors.The results show that the rotated vectors outperform the original and the sparse overcomplete vectors for interpretability and expressiveness tasks.
JinYeong Bak, Alice Oh
EMNLP2
2014 Self-disclosure topic model for classifying and analyzing Twitter conversations
abstract
Self-disclosure, the act of revealing one-self to others, is an important social be-havior that strengthens interpersonal rela-tionships and increases social support. Al-though there are many social science stud-ies of self-disclosure, they are based on manual coding of small datasets and ques-tionnaires. We conduct a computational analysis of self-disclosure with a large dataset of naturally-occurring conversa-tions, a semi-supervised machine learning algorithm, and a computational analysis of the effects of self-disclosure on subse-quent conversations. We use a longitu-dinal dataset of 17 million tweets, all of which occurred in conversations that con-sist of five or more tweets directly reply-ing to the previous tweet, and from dyads with twenty of more conversations each. We develop self-disclosure topic model (SDTM), a variant of latent Dirichlet al-location (LDA) for automatically classi-fying the level of self-disclosure for each tweet. We take the results of SDTM and analyze the effects of self-disclosure on subsequent conversations. Our model sig-nificantly outperforms several comparable methods on classifying the level of self-disclosure, and the analysis of the longitu-dinal data using SDTM uncovers signifi-cant and positive correlation between self-disclosure and conversation frequency and length. 1
JinYeong Bak, Chin-Yew Lin, Alice Oh
EMNLP1
2012 Do You Feel What I Feel? Social Aspects of Emotions in Twitter Conversations
Suin Kim, JinYeong Bak, Alice Oh
ICWSM2