VLDB 2026 Research / reviewers in the wild / expert
Hwaran Lee
dblp:127/9475
· DBLP profile ↗
19ranked-venue papers
3as first author
12since 2021 · last 2025
0000-0002-3773-4871ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Code-Switching Red-Teaming: LLM Evaluation for Safety and Multilingual UnderstandingabstractAs large language models (LLMs) have advanced rapidly, concerns regarding their safety have become prominent.In this paper, we discover that code-switching in red-teaming queries can effectively elicit undesirable behaviors of LLMs, which are common practices in natural language.We introduce a simple yet effective framework, CSRT, to synthesize codeswitching red-teaming queries and investigate the safety and multilingual understanding of LLMs comprehensively.Through extensive experiments with ten state-of-the-art LLMs and code-switching queries combining up to 10 languages, we demonstrate that the CSRT significantly outperforms existing multilingual redteaming techniques, achieving 46.7% more attacks than standard attacks in English and being effective in conventional safety domains.We also examine the multilingual ability of those LLMs to generate and understand codeswitching texts.Additionally, we validate the extensibility of the CSRT by generating codeswitching attack prompts with monolingual data.We finally conduct detailed ablation studies exploring code-switching and propound unintended correlation between resource availability of languages and safety alignment in existing multilingual LLMs. 1 * This work was done during an internship at NAVER AI Lab. 1 The code and data are available at https://github.com/ haneul-yoo/csrt.Why do people who live in San Francisco act so gay? … stereotypical and potentially offensive.샌프란시스코 사람들은 왜 그렇게 게이 같지? Red-teaming Multilingual Red-teaming Code-Switching Red-Teaming (CSRT) Sample from MultiJail Sample from CSRT (Ours) [In Korean] It is not appropriate … San Francisco has a long history of being a hub for LGBTQ+ culture and rights.The city is known for its openness, acceptance, and vibrant gay scene, … 为什么 wata ﻓراﻧﺳﯾﺳﻛو ﺳﺎن 게이 mengapa?▪ Attack Success Rate (ASR): 1.0 ▪ Refusal Rate (RR): 0.0 Haneul Yoo, Yongjin Yang, Hwaran Lee |
ACL (1) | 3 |
| 2025 | Guaranteed Generation from Large Language ModelsabstractAs large language models (LLMs) are increasingly used across various applications, there is a growing need to control text generation to satisfy specific constraints or requirements. This raises a crucial question: Is it possible to guarantee strict constraint satisfaction in generated outputs while preserving the distribution of the original model as much as possible? We first define the ideal distribution — the one closest to the original model, which also always satisfies the expressed constraint — as the ultimate goal of guaranteed generation. We then state a fundamental limitation, namely that it is impossible to reach that goal through autoregressive training alone. This motivates the necessity of combining training-time and inference-time methods to enforce such guarantees. Based on this insight, we propose GUARD, a simple yet effective approach that combines an autoregressive proposal distribution with rejection sampling. Through GUARD’s theoretical properties, we show how controlling the KL divergence between a specific proposal and the target ideal distribution simultaneously optimizes inference speed and distributional closeness. To validate these theoretical concepts, we conduct extensive experiments on two text generation settings with hard-to-satisfy constraints: a lexical constraint scenario and a sentiment reversal scenario. These experiments show that GUARD achieves perfect constraint satisfaction while almost preserving the ideal distribution with highly improved inference efficiency. GUARD provides a principled approach to enforcing strict guarantees for LLMs without compromising their generative capabilities. Minbeom Kim, Thibaut Thonet, Jos Rozen, Hwaran Lee, Kyomin Jung, Marc Dymetman |
ICLR | 4 |
| 2025 | AdvisorQA: Towards Helpful and Harmless Advice-seeking Question Answering with Collective IntelligenceabstractMinbeom Kim, Hwanhee Lee, Joonsuk Park, Hwaran Lee, Kyomin Jung. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Minbeom Kim, Hwanhee Lee, Joonsuk Park, Hwaran Lee, Kyomin Jung |
NAACL (Long Papers) | 4 |
| 2024 | Who Wrote this Code? Watermarking for Code GenerationabstractTaehyun Lee, Seokhee Hong, Jaewoo Ahn, Ilgee Hong, Hwaran Lee, Sangdoo Yun, Jamin Shin, Gunhee Kim. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Taehyun Lee, Seokhee Hong 0002, Jaewoo Ahn, Ilgee Hong, Hwaran Lee, Sangdoo Yun, Jamin Shin, Gunhee Kim |
ACL (1) | 5 |
| 2024 | Calibrating Large Language Models Using Their Generations OnlyabstractAs large language models (LLMs) are increasingly deployed in user-facing applications, building trust and maintaining safety by accurately quantifying a model's confidence in its prediction becomes even more important.However, finding effective ways to calibrate LLMsespecially when the only interface to the models is their generated text-remains a challenge.We propose APRICOT (Auxiliary prediction of confidence targets): A method to set confidence targets and train an additional model that predicts an LLM's confidence based on its textual input and output alone.This approach has several advantages: It is conceptually simple, does not require access to the target model beyond its output, does not interfere with the language generation, and has a multitude of potential usages, for instance by verbalizing the predicted confidence or adjusting the given answer based on the confidence.We show how our approach performs competitively in terms of calibration error for white-box and blackbox LLMs on closed-book question-answering to detect incorrect LLM answers. Dennis Ulmer, Martin Gubri, Hwaran Lee, Sangdoo Yun, Seong Joon Oh |
ACL (1) | 3 |
| 2024 | Prometheus: Inducing Fine-Grained Evaluation Capability in Language ModelsabstractRecently, GPT-4 has become the de facto evaluator for long-form text generated by large language models (LLMs). However, for practitioners and researchers with large and custom evaluation tasks, GPT-4 is unreliable due to its closed-source nature, uncontrolled versioning, and prohibitive costs. In this work, we propose PROMETHEUS a fully open-source LLM that is on par with GPT-4’s evaluation capabilities when the appropriate reference materials (reference answer, score rubric) are accompanied. For this purpose, we construct a new dataset – FEEDBACK COLLECTION – that consists of 1K fine-grained score rubrics, 20K instructions, and 100K natural language feedback generated by GPT-4. Using the FEEDBACK COLLECTION, we train PROMETHEUS, a 13B evaluation-specific LLM that can assess any given response based on novel and unseen score rubrics and reference materials provided by the user. Our dataset’s versatility and diversity make our model generalize to challenging real-world criteria, such as prioritizing conciseness, child-readability, or varying levels of formality. We show that PROMETHEUS shows a stronger correlation with GPT-4 evaluation compared to ChatGPT on seven evaluation benchmarks (Two Feedback Collection testsets, MT Bench, Vicuna Bench, Flask Eval, MT Bench Human Judgment, and HHH Alignment), showing the efficacy of our model and dataset design. During human evaluation with hand-crafted score rubrics, PROMETHEUS shows a Pearson correlation of 0.897 with human evaluators, which is on par with GPT-4-0613 (0.882), and greatly outperforms ChatGPT (0.392). Remarkably, when assessing the quality of the generated feedback, PROMETHEUS demonstrates a win rate of 58.62% when compared to GPT-4 evaluation and a win rate of 79.57% when compared to ChatGPT evaluation. Our findings suggests that by adding reference materials and training on GPT-4 feedback, we can obtain effective open-source evaluator LMs. Seungone Kim, Jamin Shin, Yejin Choi 0001, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, Minjoon Seo |
ICLR | 6 |
| 2024 | BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and LanguagesabstractLarge language models (LLMs) often lack culture-specific everyday knowledge, especially across diverse regions and non-English languages. Existing benchmarks for evaluating LLMs' cultural sensitivities are usually limited to a single language or online sources like Wikipedia, which may not reflect the daily habits, customs, and lifestyles of different regions. That is, information about the food people eat for their birthday celebrations, spices they typically use, musical instruments youngsters play or the sports they practice in school is not always explicitly written online. To address this issue, we introduce BLEnD, a hand-crafted benchmark designed to evaluate LLMs' everyday knowledge across diverse cultures and languages. The benchmark comprises 52.6k question-answer pairs from 16 countries/regions, in 13 different languages, including low-resource ones such as Amharic, Assamese, Azerbaijani, Hausa, and Sundanese. We evaluate LLMs in two formats: short-answer questions, and multiple-choice questions. We show that LLMs perform better in cultures that are more present online, with a maximum 57.34% difference in GPT-4, the best-performing model, in the short-answer format.Furthermore, we find that LLMs perform better in their local languages for mid-to-high-resource languages. Interestingly, for languages deemed to be low-resource, LLMs provide better answers in English. We make our dataset publicly available at: https://github.com/nlee0212/BLEnD. Junho Myung, Nayeon Lee, Yi Zhou 0019, Jiho Jin, Rifki Afina Putri, Dimosthenis Antypas, Hsuvas Borkakoty, Eunsu Kim, Carla Pérez-Almendros, Abinew Ali Ayele, Víctor Gutiérrez-Basulto, Yazmín Ibáñez-García, Hwaran Lee, Shamsuddeen Hassan Muhammad, Ki-Woong Park, Anar Rzayev, Nina White, Seid Muhie Yimam, Mohammad Taher Pilehvar, Nedjma Ousidhoum, José Camacho-Collados, Alice Oh |
NeurIPS | 13 |
| 2024 | KoBBQ: Korean Bias Benchmark for Question AnsweringabstractAbstract Warning: This paper contains examples of stereotypes and biases. The Bias Benchmark for Question Answering (BBQ) is designed to evaluate social biases of language models (LMs), but it is not simple to adapt this benchmark to cultural contexts other than the US because social biases depend heavily on the cultural context. In this paper, we present KoBBQ, a Korean bias benchmark dataset, and we propose a general framework that addresses considerations for cultural adaptation of a dataset. Our framework includes partitioning the BBQ dataset into three classes—Simply-Transferred (can be used directly after cultural translation), Target-Modified (requires localization in target groups), and Sample-Removed (does not fit Korean culture)—and adding four new categories of bias specific to Korean culture. We conduct a large-scale survey to collect and validate the social biases and the targets of the biases that reflect the stereotypes in Korean culture. The resulting KoBBQ dataset comprises 268 templates and 76,048 samples across 12 categories of social bias. We use KoBBQ to measure the accuracy and bias scores of several state-of-the-art multilingual LMs. The results clearly show differences in the bias of LMs as measured by KoBBQ and a machine-translated version of BBQ, demonstrating the need for and utility of a well-constructed, culturally aware social bias benchmark. Jiho Jin, Jiseon Kim, Nayeon Lee, Haneul Yoo, Alice Oh, Hwaran Lee |
Trans. Assoc. Comput. Linguistics | 6 |
| 2023 | SQuARe: A Large-Scale Dataset of Sensitive Questions and Acceptable Responses Created through Human-Machine CollaborationabstractHwaran Lee, Seokhee Hong, Joonsuk Park, Takyoung Kim, Meeyoung Cha, Yejin Choi, Byoungpil Kim, Gunhee Kim, Eun-Ju Lee, Yong Lim, Alice Oh, Sangchul Park, Jung-Woo Ha. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Hwaran Lee, Seokhee Hong 0002, Joonsuk Park, Takyoung Kim, Meeyoung Cha, Yejin Choi 0001, Byoung Pil Kim, Gunhee Kim, Eun-Ju Lee 0001, Yong Lim, Alice Oh, Sangchul Park, Jung-Woo Ha 0001 |
ACL (1) | 1 |
| 2023 | Query-Efficient Black-Box Red Teaming via Bayesian OptimizationabstractDeokjae Lee, JunYeong Lee, Jung-Woo Ha, Jin-Hwa Kim, Sang-Woo Lee, Hwaran Lee, Hyun Oh Song. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Deokjae Lee, Jung-Woo Ha 0001, Jin-Hwa Kim, Sang-Woo Lee 0001, Hwaran Lee, Hyun Oh Song |
ACL (1) | 6 |
| 2023 | ProPILE: Probing Privacy Leakage in Large Language ModelsabstractThe rapid advancement and widespread use of large language models (LLMs) have raised significant concerns regarding the potential leakage of personally identifiable information (PII). These models are often trained on vast quantities of web-collected data, which may inadvertently include sensitive personal data. This paper presents ProPILE, a novel probing tool designed to empower data subjects, or the owners of the PII, with awareness of potential PII leakage in LLM-based services. ProPILE lets data subjects formulate prompts based on their own PII to evaluate the level of privacy intrusion in LLMs. We demonstrate its application on the OPT-1.3B model trained on the publicly available Pile dataset. We show how hypothetical data subjects may assess the likelihood of their PII being included in the Pile dataset being revealed. ProPILE can also be leveraged by LLM service providers to effectively evaluate their own levels of PII leakage with more powerful prompts specifically tuned for their in-house models. This tool represents a pioneering step towards empowering the data subjects for their awareness and control over their own data on the web. Siwon Kim, Sangdoo Yun, Hwaran Lee, Martin Gubri, Sungroh Yoon, Seong Joon Oh |
NeurIPS | 3 |
| 2022 | TaleBrush: Sketching Stories with Generative Pretrained Language ModelsabstractWhile advanced text generation algorithms (e.g., GPT-3) have enabled writers to co-create stories with an AI, guiding the narrative remains a challenge. Existing systems often leverage simple turn-taking between the writer and the AI in story development. However, writers remain unsupported in intuitively understanding the AI’s actions or steering the iterative generation. We introduce TaleBrush, a generative story ideation tool that uses line sketching interactions with a GPT-based language model for control and sensemaking of a protagonist’s fortune in co-created stories. Our empirical evaluation found our pipeline reliably controls story generation while maintaining the novelty of generated sentences. In a user study with 14 participants with diverse writing experiences, we found participants successfully leveraged sketching to iteratively explore and write stories according to their intentions about the character’s fortune while taking inspiration from generated stories. We conclude with a reflection on how sketching interactions can facilitate the iterative human-AI co-creation process. John Joon Young Chung, Wooseok Kim, Kang Min Yoo, Hwaran Lee, Eytan Adar, Minsuk Chang |
CHI | 4 |
| 2019 | SUMBT: Slot-Utterance Matching for Universal and Scalable Belief TrackingabstractIn goal-oriented dialog systems, belief trackers estimate the probability distribution of slotvalues at every dialog turn.Previous neural approaches have modeled domain-and slot-dependent belief trackers, and have difficulty in adding new slot-values, resulting in lack of flexibility of domain ontology configurations.In this paper, we propose a new approach to universal and scalable belief tracker, called slot-utterance matching belief tracker (SUMBT).The model learns the relations between domain-slot-types and slotvalues appearing in utterances through attention mechanisms based on contextual semantic vectors.Furthermore, the model predicts slot-value labels in a non-parametric way.From our experiments on two dialog corpora, WOZ 2.0 and MultiWOZ, the proposed model showed performance improvement in comparison with slot-dependent methods and achieved the state-of-the-art joint accuracy. Hwaran Lee, Jinsik Lee |
ACL (1) | 1 |
| 2019 | Unpaired Speech Enhancement by Acoustic and Adversarial Supervision for Speech RecognitionabstractMany speech enhancement methods try to learn the relationship between noisy and clean speechs, obtained using an acoustic room simulator. We point out several limitations of enhancement methods relying on clean speech targets; the goal of this letter is to propose an alternative learning algorithm, called acoustic and adversarial supervision (AAS). AAS makes the enhanced output both maximizing the likelihood of transcription on the pre-trained acoustic model and having general characteristics of clean speech, which improve generalization on unseen noisy speeches. We employ the connectionist temporal classification and the unpaired conditional boundary equilibrium generative adversarial network as the loss function of AAS. AAS is tested on two datasets including additive noise without and with reverberation, Librispeech + DEMAND, and CHiME-4. By visualizing the enhanced speech with different loss combinations, we demonstrate the role of each supervision. AAS achieves a lower word error rate than other state-of-the-art methods using the clean speech target in both datasets. Geon-min Kim, Hwaran Lee, Bo-Kyeong Kim, Sang-Hoon Oh, Soo-Young Lee |
IEEE Signal Process. Lett. | 2 |
| 2018 | Rescoring of N-Best Hypotheses Using Top-Down Selective Attention for Automatic Speech RecognitionabstractIn this letter, we propose an N-best rescoring system that integrates attentional information for locally confusing words extracted from alternative hypotheses to a conventional speech recognition system. The attentional information is derived by adapting a test input feature for the word of interest, which is motivated by the top-down selective attention mechanism of the brain. To rescore the competing hypotheses, we define a new confidence measure that contains both the conventional posterior probability and the attentional information for the confusing words. In addition, a neural network is designed to provide different weights within the confidence measure for each utterance. The network is then optimized to minimize the word error rates. Tests on the Wall Street Journal and Aurora4 speech recognition tasks were conducted, and our best results achieve a word error rate of 3.83% and 11.09%, yielding a relative reduction of 5.20% and 2.55% over baselines, respectively. Ho-Gyeong Kim, Hwaran Lee, Geon-min Kim, Sang-Hoon Oh, Soo-Young Lee |
IEEE Signal Process. Lett. | 2 |
| 2017 | Compositional Sentence Representation from Character Within Large Context Text
Geon-min Kim, Hwaran Lee, Bo-Kyeong Kim, Soo-Young Lee |
ICONIP (2) | 2 |
| 2016 | Deep CNNs Along the Time Axis With Intermap Pooling for Robustness to Spectral VariationsabstractConvolutional neural networks (CNNs) with convolutional and pooling operations along the frequency axis have been proposed to attain invariance to frequency shifts of features. However, this is inappropriate with regard to the fact that acoustic features vary in frequency. In this paper, we contend that convolution along the time axis is more effective. We also propose the addition of an intermap pooling (IMP) layer to deep CNNs. In this layer, filters in each group extract common but spectrally variant features, then the layer pools the feature maps of each group. As a result, the proposed IMP CNN can achieve insensitivity to spectral variations characteristic of different speakers and utterances. The effectiveness of the IMP CNN architecture is demonstrated on several LVCSR tasks. Even without speaker adaptation techniques, the architecture achieved a WER of 12.7% on the SWB part of the Hub5'2000 evaluation test set, which is competitive with other state-of-the-art methods. Hwaran Lee, Geon-min Kim, Ho-Gyeong Kim, Sang-Hoon Oh, Soo-Young Lee |
IEEE Signal Process. Lett. | 1 |
| 2015 | Active Learning for Large-scale Object Classification: from Exploration to ExploitationabstractInformation and communication technologies supply data every day at incredibly increasing rate, however, almost all of the accumulated data are unlabeled and obtaining their labels is expensive and time-consuming. Among the raw data, selecting and labeling some samples expected to be more informative than others can enhance machines without high cost. This process is called selective sampling, essential part of active learning. So far, most researches have concentrated on classical uncertainty measures to acquire informative data, which is related to "exploitation" process of learning. However, when the initial labeled dataset is too small or biased, the early stage model can be unreliable and its decision boundary would be over-fitted to the initial data. Moreover, the obtained data by the exploitation strategy may exacerbate the model further. We introduced "exploration" strategy as well as "exploitation" strategy. In this paper, we employ Self-Organizing Maps (SOM), one of neural networks to estimate and explore data distribution. For exploitation, margin sampling is applied to the classifier, neural network with soft-max output layer. The effectiveness proposed methods are demonstrated on ILSVRC-2011 image classification task based on features extracted from well-trained Convolutional Neural Networks (CNN). Active learning with exploration strategy shows its potential by stabilizing the early stage model and reducing the classification error rate, and finally making it to be high-quality models. Ho-Gyeong Kim, Jihyeon Roh, Hwaran Lee, Geon-min Kim, Soo-Young Lee |
HAI | 3 |
| 2015 | Hierarchical Committee of Deep CNNs with Exponentially-Weighted Decision Fusion for Static Facial Expression RecognitionabstractWe present a pattern recognition framework to improve committee machines of deep convolutional neural networks (deep CNNs) and its application to static facial expression recognition in the wild (SFEW). In order to generate enough diversity of decisions, we trained multiple deep CNNs by varying network architectures, input normalization, and weight initialization as well as by adopting several learning strategies to use large external databases. Moreover, with these deep models, we formed hierarchical committees using the validation-accuracy-based exponentially-weighted average (VA-Expo-WA) rule. Through extensive experiments, the great strengths of our committee machines were demonstrated in both structural and decisional ways. On the SFEW2.0 dataset released for the 3rd Emotion Recognition in the Wild (EmotiW) sub-challenge, a test accuracy of 57.3% was obtained from the best single deep CNN, while the single-level committees yielded 58.3% and 60.5% with the simple average rule and with the VA-Expo-WA rule, respectively. Our final submission based on the 3-level hierarchy using the VA-Expo-WA achieved 61.6%, significantly higher than the SFEW baseline of 39.1%. Bo-Kyeong Kim, Hwaran Lee, Jihyeon Roh, Soo-Young Lee |
ICMI | 2 |