VLDB 2026 Research / reviewers in the wild / expert
Joongbo Shin
dblp:207/7602
· DBLP profile ↗
13ranked-venue papers
1as first author
9since 2021 · last 2025
0009-0004-8957-4867ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FaithUn: Toward Faithful Forgetting in Language Models by Investigating the Interconnectedness of KnowledgeabstractVarious studies have attempted to remove sensitive or private knowledge from a language model to prevent its unauthorized exposure.However, prior studies have overlooked the inherent complexity and interconnectedness of knowledge, which requires careful examination.To resolve this problem, we first define a new concept called superficial unlearning, which refers to the phenomenon where an unlearning method either fails to erase the interconnected knowledge it should remove or unintentionally erases irrelevant knowledge.Based on the definition, we introduce a novel benchmark, FAITHUN, to analyze and evaluate the faithfulness of unlearning in real-world knowledge QA settings.Furthermore, we propose a novel unlearning method, KLUE, which updates only knowledge-related neurons to achieve faithful unlearning.KLUE leverages a regularized explainability method to localize contextual knowledge neurons, updating only these neurons using carefully selected unforgotten samples.Experimental results demonstrate that existing unlearning methods fail to ensure faithful unlearning, while our method shows significant effectiveness in real-world QA unlearning. Nakyeong Yang, Seunghyun Yoon 0002, Joongbo Shin, Kyomin Jung |
EMNLP | 4 |
| 2024 | A New Framework for Evaluating Faithfulness of Video Moment Retrieval against Multiple DistractorsabstractWith the explosion of multimedia content, video moment retrieval (VMR), which aims to detect a video moment that matches a given text query from a video, has been studied intensively as a critical problem. However, the existing VMR framework evaluates video moment retrieval performance, assuming that a video is given, which may not reveal whether the models exhibit overconfidence in the falsely given video. In this paper, we propose the MVMR (Massive Videos Moment Retrieval for Faithfulness Evaluation) task that aims to retrieve video moments within a massive video set, including multiple distractors, to evaluate the faithfulness of VMR models. For this task, we suggest an automated massive video pool construction framework to categorize negative (distractors) and positive (false-negative) video sets using textual and visual semantic distance verification methods. We extend existing VMR datasets using these methods and newly construct three practical MVMR datasets. To solve the task, we further propose a strong informative sample-weighted learning method, CroCs, which employs two contrastive learning mechanisms: (1) weakly-supervised potential negative learning and (2) cross-directional hard-negative learning. Experimental results on the MVMR datasets reveal that existing VMR models are easily distracted by the misinformation (distractors), whereas our model shows significantly robust performance, demonstrating that CroCs is essential to distinguishing positive moments against distractors. Nakyeong Yang, Seunghyun Yoon 0002, Joongbo Shin, Kyomin Jung |
CIKM | 4 |
| 2024 | Exploring the Use of Natural Language Descriptions of Intents for Large Language Models in Zero-shot Intent ClassificationabstractTaesuk Hong, Youbin Ahn, Dongkyu Lee, Joongbo Shin, Seungpil Won, Janghoon Han, Stanley Jungkyu Choi, Jungyun Seo. Proceedings of the 25th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2024. Taesuk Hong, Youbin Ahn, Joongbo Shin, Seungpil Won, Janghoon Han, Stanley Jungkyu Choi, Jungyun Seo |
SIGDIAL | 4 |
| 2023 | BREAK: Breaking the Dialogue State Tracking Barrier with Beam Search and Re-rankingabstractDespite the recent advances in dialogue state tracking (DST), the joint goal accuracy (JGA) of the existing methods on MultiWOZ 2.1 still remains merely 60%.In our preliminary error analysis, we find that beam search produces a pool of candidates that is likely to include the correct dialogue state.Motivated by this observation, we introduce a novel framework, called BREAK (Beam search and RE-rAnKing), that achieves outstanding performance on DST.Our proposed method performs DST in two stages: (i) generating k-best dialogue state candidates with beam search and (ii) re-ranking the candidates to select the correct dialogue state.This simple yet powerful framework shows state-of-the-art performance on all versions of MultiWOZ and M2M datasets.Most notably, we push the joint goal accuracy to 80-90% on MultiWOZ 2.1-2.4,which is an improvement of 23.6%, 26.3%, 21.7%, and 10.8% over the previous best-performing models, respectively. Seungpil Won, Heeyoung Kwak, Joongbo Shin, Janghoon Han, Kyomin Jung |
ACL (1) | 3 |
| 2023 | Guess the Instruction! Flipped Learning Makes Language Models Stronger Zero-Shot Learners
Seonghyeon Ye, Doyoung Kim 0001, Joel Jang, Joongbo Shin, Minjoon Seo |
ICLR | 4 |
| 2022 | TemporalWiki: A Lifelong Benchmark for Training and Evaluating Ever-Evolving Language ModelsabstractLanguage Models (LMs) become outdated as the world changes; they often fail to perform tasks requiring recent factual information which was absent or different during training, a phenomenon called temporal misalignment.This is especially a challenging problem because the research community still lacks a coherent dataset for assessing the adaptability of LMs to frequently-updated knowledge corpus such as Wikipedia.To this end, we introduce TEMPORALWIKI, a lifelong benchmark for ever-evolving LMs that utilizes the difference between consecutive snapshots of English Wikipedia and English Wikidata for training and evaluation, respectively.The benchmark hence allows researchers to periodically track an LM's ability to retain previous knowledge and acquire updated/new knowledge at each point in time.We also find that training an LM on the diff data through continual learning methods achieves similar or better perplexity than on the entire snapshot in our benchmark with 12 times less computational cost, which verifies that factual knowledge in LMs can be safely updated with minimal training data via continual learning.The dataset and the code is made available at this link. What is the most dominant COVID-19 variant?What is the most dominant COVID-19 variant?Difference Joel Jang, Seonghyeon Ye, Sohee Yang, Joongbo Shin, Janghoon Han, Gyeonghun Kim, Minjoon Seo |
EMNLP | 5 |
| 2022 | Towards Continual Knowledge Learning of Language Models
Joel Jang, Seonghyeon Ye, Sohee Yang, Joongbo Shin, Janghoon Han, Gyeonghun Kim, Stanley Jungkyu Choi, Minjoon Seo |
ICLR | 4 |
| 2021 | Bidirectional Variational Inference for Non-Autoregressive Text-to-Speech
Yoonhyung Lee, Joongbo Shin, Kyomin Jung |
ICLR | 2 |
| 2021 | KPQA: A Metric for Generative Question Answering Using Keyphrase WeightsabstractHwanhee Lee, Seunghyun Yoon, Franck Dernoncourt, Doo Soon Kim, Trung Bui, Joongbo Shin, Kyomin Jung. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Hwanhee Lee, Seunghyun Yoon 0002, Franck Dernoncourt, Doo Soon Kim, Trung Bui, Joongbo Shin, Kyomin Jung |
NAACL-HLT | 6 |
| 2020 | Fast and Accurate Deep Bidirectional Language Representations for Unsupervised LearningabstractEven though BERT has achieved successful performance improvements in various supervised learning tasks, BERT is still limited by repetitive inferences on unsupervised tasks for the computation of contextual language representations.To resolve this limitation, we propose a novel deep bidirectional language model called a Transformer-based Text Autoencoder (T-TA).The T-TA computes contextual language representations without repetition and displays the benefits of a deep bidirectional architecture, such as that of BERT.In computation time experiments in a CPU environment, the proposed T-TA performs over six times faster than the BERT-like model on a reranking task and twelve times faster on a semantic similarity task.Furthermore, the T-TA shows competitive or even better accuracies than those of BERT on the above tasks.Code is available at https://github.com/joongbo/tta. Joongbo Shin, Yoonhyung Lee, Seunghyun Yoon 0002, Kyomin Jung |
ACL | 1 |
| 2019 | Detecting Incongruity between News Headline and Body Text via a Deep Hierarchical EncoderabstractSome news headlines mislead readers with overrated or false information, and identifying them in advance will better assist readers in choosing proper news stories to consume. This research introduces million-scale pairs of news headline and body text dataset with incongruity label, which can uniquely be utilized for detecting news stories with misleading headlines. On this dataset, we develop two neural networks with hierarchical architectures that model a complex textual representation of news articles and measure the incongruity between the headline and the body text. We also present a data augmentation method that dramatically reduces the text input size a model handles by independently investigating each paragraph of news stories, which further boosts the performance. Our experiments and qualitative evaluations demonstrate that the proposed methods outperform existing approaches and efficiently detect news stories with misleading headlines in the real world. Seunghyun Yoon 0002, Kunwoo Park, Joongbo Shin, Hongjun Lim, Seungpil Won, Meeyoung Cha, Kyomin Jung |
AAAI | 3 |
| 2019 | Improving Neural Question Generation Using Answer SeparationabstractNeural question generation (NQG) is the task of generating a question from a given passage with deep neural networks. Previous NQG models suffer from a problem that a significant proportion of the generated questions include words in the question target, resulting in the generation of unintended questions. In this paper, we propose answer-separated seq2seq, which better utilizes the information from both the passage and the target answer. By replacing the target answer in the original passage with a special token, our model learns to identify which interrogative word should be used. We also propose a new module termed keyword-net, which helps the model better capture the key information in the target answer and generate an appropriate question. Experimental results demonstrate that our answer separation method significantly reduces the number of improper questions which include answers. Consequently, our model significantly outperforms previous state-of-the-art NQG models. Yanghoon Kim, Hwanhee Lee, Joongbo Shin, Kyomin Jung |
AAAI | 3 |
| 2018 | Learning to Rank Question-Answer Pairs Using Hierarchical Recurrent Encoder with Latent Topic ClusteringabstractSeunghyun Yoon, Joongbo Shin, Kyomin Jung. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Seunghyun Yoon 0002, Joongbo Shin, Kyomin Jung |
NAACL-HLT | 2 |