VLDB 2026 Research / reviewers in the wild / expert
Seungtaek Choi
dblp:218/7548
· DBLP profile ↗
25ranked-venue papers
4as first author
17since 2021 · last 2025
0000-0003-3570-0907ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 3 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Rethinking the Training Paradigm of Discrete Token-Based Multimodal LLMs: An Analysis of Text-Centric BiasabstractDiscrete token-based multimodal large language models (MLLMs), such as AnyGPT and MIO, integrate diverse modalities into an autoregressive framework by discretizing modality inputs into tokens compatible with language models. Unlike encoder-based approaches, such as LLaVA and Flamingo, which utilize pretrained modality-specific encoders, discrete token-based MLLMs simultaneously learn modality token representations and their alignment with the language, yet are exclusively trained on modality-text paired datasets without additional unimodal training. We identify a structural limitation inherent in this training paradigm, termed text-centric bias, defined as an over-reliance on the textual context that restricts intrinsic modality understanding. To systematically analyze the existence of this bias, we propose an analytical framework involving external perplexity-based and internal neuron-level analyses. Furthermore, to verify whether the bias originates from the paired-only training paradigm, we introduce an analytical methodology named Monotune, which is a simple unimodal training stage. Our analyses demonstrate that minimal exposure to unimodal data effectively mitigates text-centric bias, providing empirical evidence that the bias is fundamentally induced by the paired-only training strategy. Through comprehensive downstream task evaluations, we further reveal that this structural bias meaningfully affects real-world multimodal task performance, particularly under limited textual contexts. Our findings highlight a fundamental limitation in current discrete token-based MLLM training paradigms and suggest directions for future multimodal training strategies. Our code and experiments are available at https://github.com/41312432/Monotune Wansik Jo, Jooyeong Na, Soyeon Hong, Seungtaek Choi, Hyunsouk Cho |
CIKM | 4 |
| 2025 | Towards Fully-Automated Materials Discovery via Large-Scale Synthesis Dataset and Expert-Level LLM-as-a-JudgeabstractMaterials synthesis remains a critical bottleneck in developing innovations for energy storage, catalysis, electronics, and biomedical devices. Current synthesis design relies heavily on empirical trial-and-error methods guided by expert intuition, limiting the pace of materials discovery. To address this challenge, we present AlchemyBench, a comprehensive benchmark built upon a curated dataset of 17,667 expert-verified synthesis recipes from open-access literature. Heegyu Kim, Taeyang Jeon, Seungtaek Choi, Jihoon Hong, Dongwon Jeon, Ga-Yeon Baek, Gyeong-Won Kwak, Jisu Bae, Yoon-Seo Kim, Seon-Jin Choi, Sung Beom Cho, Hyunsouk Cho |
CIKM | 3 |
| 2025 | FLEX: Expert-level False-Less EXecution Metric for Text-to-SQL BenchmarkabstractHeegyu Kim, Jeon Taeyang, SeungHwan Choi, Seungtaek Choi, Hyunsouk Cho. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Heegyu Kim, Taeyang Jeon, Seunghwan Choi, Seungtaek Choi, Hyunsouk Cho |
NAACL (Long Papers) | 4 |
| 2025 | ScoreCL: augmentation-adaptive contrastive learning via score-matching function
Soonwoo Kwon, Hyojun Go, Yunsung Lee, Seungtaek Choi, Hyun-Gyoon Kim |
Mach. Learn. | 5 |
| 2025 | Overcoming Source Object Grounding for Semantic Image EditingabstractAbstract Recent diffusion models have demonstrated remarkable capabilities in text-to-image generation. However, their stochastic denoising process often causes semantic image editing (SIE) models to misapply textual instructions. That is, models often leave the source object unchanged or erroneously alter the background. We refer to this challenge as source object grounding. To address this challenge, we introduce R-SIE, a region-wise SIE framework. During the inference, R-SIE models noise separately for distinct image regions, enabling precise control over the transformed areas. To reinforce the inference, we devise an automatic pipeline leveraging bounding boxes to generate unambiguous training data. Additionally, we propose two region-focused metrics, CLIP-Region Class (CLIP-RC) and CLIP-Global Context (CLIP-GC), to independently assess how well the source object is edited and the background is preserved, respectively. Experimental results demonstrate that region-wise diffusion improves existing baselines, and our data generation pipeline further enhances these improvements.1 YeonJoon Jung, Seungtaek Choi, Seung-won Hwang |
Trans. Assoc. Comput. Linguistics | 2 |
| 2024 | Multi-Architecture Multi-Expert Diffusion ModelsabstractIn this paper, we address the performance degradation of efficient diffusion models by introducing Multi-architecturE Multi-Expert diffusion models (MEME). We identify the need for tailored operations at different time-steps in diffusion processes and leverage this insight to create compact yet high-performing models. MEME assigns distinct architectures to different time-step intervals, balancing convolution and self-attention operations based on observed frequency characteristics. We also introduce a soft interval assignment strategy for comprehensive training. Empirically, MEME operates 3.3 times faster than baselines while improving image generation quality (FID scores) by 0.62 (FFHQ) and 0.37 (CelebA). Though we validate the effectiveness of assigning more optimal architecture per time-step, where efficient models outperform the larger models, we argue that MEME opens a new design choice for diffusion models that can be easily applied in other scenarios, such as large multi-expert models. Yunsung Lee, Hyojun Go, Myeongho Jeong, Shinhyeok Oh, Seungtaek Choi |
AAAI | 6 |
| 2024 | Interventional Speech Noise Injection for ASR Generalizable Spoken Language UnderstandingabstractRecently, pre-trained language models (PLMs) have been increasingly adopted in spoken language understanding (SLU).However, automatic speech recognition (ASR) systems frequently produce inaccurate transcriptions, leading to noisy inputs for SLU models, which can significantly degrade their performance.To address this, our objective is to train SLU models to withstand ASR errors by exposing them to noises commonly observed in ASR systems, referred to as ASR-plausible noises.Speech noise injection (SNI) methods have pursued this objective by introducing ASR-plausible noises, but we argue that these methods are inherently biased towards specific ASR systems, or ASR-specific noises.In this work, we propose a novel and less biased augmentation method of introducing the noises that are plausible to any ASR system, by cutting off the non-causal effect of noises.Experimental results and analyses demonstrate the effectiveness of our proposed methods in enhancing the robustness and generalizability of SLU models against unseen ASR systems by introducing more diverse and plausible ASR noises in advance. YeonJoon Jung, Jaeseong Lee 0002, Seungtaek Choi, Dohyeon Lee, Seung-won Hwang |
EMNLP | 3 |
| 2023 | On Complementarity Objectives for Hybrid RetrievalabstractDense retrieval has shown promising results in various information retrieval tasks, and hybrid retrieval, combined with the strength of sparse retrieval, has also been actively studied.A key challenge in hybrid retrieval is to make sparse and dense complementary to each other.Existing models have focused on dense models to capture "residual" features neglected in the sparse models.Our key distinction is to show how this notion of residual complementarity is limited, and propose a new objective, denoted as RoC (Ratio of Complementarity), which captures a fuller notion of complementarity.We propose a two-level orthogonality designed to improve RoC, then show that the improved RoC of our model, in turn, improves the performance of hybrid retrieval.Our method outperforms all state-of-the-art methods on three representative IR benchmarks: MSMARCO-Passage, Natural Questions, and TREC Ro-bust04, with statistical significance.Our finding is also consistent in various adversarial settings. Dohyeon Lee, Seung-won Hwang, Kyungjae Lee 0002, Seungtaek Choi, Sunghyun Park 0005 |
ACL (1) | 4 |
| 2023 | Towards Practical Plug-and-Play Diffusion ModelsabstractDiffusion-based generative models have achieved remarkable success in image generation. Their guidance formulation allows an external model to plug-and-play control the generation process for various tasks without finetuning the diffusion model. However, the direct use of publicly available off-the-shelf models for guidance fails due to their poor performance on noisy inputs. For that, the existing practice is to fine-tune the guidance models with labeled data corrupted with noises. In this paper, we argue that this practice has limitations in two aspects: (1) performing on inputs with extremely various noises is too hard for a single guidance model; (2) collecting labeled datasets hinders scaling up for various tasks. To tackle the limitations, we propose a novel strategy that leverages multiple experts where each expert is specialized in a particular noise range and guides the reverse process of the diffusion at its corresponding timesteps. However, as it is infeasible to manage multiple networks and utilize labeled data, we present a practical guidance framework termed Practical Plug-And-Play (PPAP), which leverages parameter-efficient fine-tuning and data-free knowledge transfer. We exhaustively conduct ImageNet class conditional generation experiments to show that our method can successfully guide diffusion with small trainable parameters and no labeled data. Finally, we show that image classifiers, depth estimators, and semantic segmentation models can guide publicly available GLIDE through our framework in a plug-and-play manner. Our code is available at https://github.com/riiid/PPAP. Hyojun Go, Yunsung Lee, Myeongho Jeong, Hyun Seung Lee, Seungtaek Choi |
CVPR | 7 |
| 2023 | Addressing Cold Start Problem for End-to-end Automatic Speech Scoring
Jungbae Park, Seungtaek Choi |
INTERSPEECH | 2 |
| 2023 | Addressing Negative Transfer in Diffusion ModelsabstractDiffusion-based generative models have achieved remarkable success in various domains. It trains a shared model on denoising tasks that encompass different noise levels simultaneously, representing a form of multi-task learning (MTL). However, analyzing and improving diffusion models from an MTL perspective remains under-explored. In particular, MTL can sometimes lead to the well-known phenomenon of $\textit{negative transfer}$, which results in the performance degradation of certain tasks due to conflicts between tasks. In this paper, we first aim to analyze diffusion training from an MTL standpoint, presenting two key observations: $\textbf{(O1)}$ the task affinity between denoising tasks diminishes as the gap between noise levels widens, and $\textbf{(O2)}$ negative transfer can arise even in diffusion training. Building upon these observations, we aim to enhance diffusion training by mitigating negative transfer. To achieve this, we propose leveraging existing MTL methods, but the presence of a huge number of denoising tasks makes this computationally expensive to calculate the necessary per-task loss or gradient. To address this challenge, we propose clustering the denoising tasks into small task clusters and applying MTL methods to them. Specifically, based on $\textbf{(O2)}$, we employ interval clustering to enforce temporal proximity among denoising tasks within clusters. We show that interval clustering can be solved using dynamic programming, utilizing signal-to-noise ratio, timestep, and task affinity for clustering objectives. Through this, our approach addresses the issue of negative transfer in diffusion models by allowing for efficient computation of MTL methods. We validate the efficacy of proposed clustering and its integration with MTL methods through various experiments, demonstrating 1) improved generation quality and 2) faster training convergence of diffusion models. Our project page is available at https://gohyojun15.github.io/ANT_diffusion/. Hyojun Go, Yunsung Lee, Shinhyeok Oh, Hyeongdon Moon, Seungtaek Choi |
NeurIPS | 7 |
| 2022 | C2L: Causally Contrastive Learning for Robust Text ClassificationabstractDespite the super-human accuracy of recent deep models in NLP tasks, their robustness is reportedly limited due to their reliance on spurious patterns. We thus aim to leverage contrastive learning and counterfactual augmentation for robustness. For augmentation, existing work either requires humans to add counterfactuals to the dataset or machines to automatically matches near-counterfactuals already in the dataset. Unlike existing augmentation is affected by spurious correlations, ours, by synthesizing “a set” of counterfactuals, and making a collective decision on the distribution of predictions on this set, can robustly supervise the causality of each term. Our empirical results show that our approach, by collective decisions, is less sensitive to task model bias of attribution-based synthesis, and thus achieves significant improvements, in diverse dimensions: 1) counterfactual robustness, 2) cross-domain generalization, and 3) generalization from scarce data. Seungtaek Choi, Myeongho Jeong, Hojae Han, Seung-won Hwang |
AAAI | 1 |
| 2022 | Towards Compositional Generalization in Code SearchabstractWe study compositional generalization, which aims to generalize on unseen combinations of seen structural elements, for code search.Unlike existing approaches of partially pursuing this goal, we study how to extract structural elements, which we name a template that directly targets compositional generalization.Thus we propose CTBERT, or Code Template BERT, representing codes using automatically extracted templates as building blocks.We empirically validate CTBERT on two public code search benchmarks, AdvTest and CSN.Further, we show that templates are complementary to data flow graphs in GraphCodeBERT, by enhancing structural context around variables. Hojae Han, Seung-won Hwang, Nan Duan 0001, Seungtaek Choi |
EMNLP | 5 |
| 2022 | Evaluating the Knowledge Dependency of QuestionsabstractHyeongdon Moon, Yoonseok Yang, Hangyeol Yu, Seunghyun Lee, Myeongho Jeong, Juneyoung Park, Jamin Shin, Minsam Kim, Seungtaek Choi. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Hyeongdon Moon, Yoonseok Yang, Hangyeol Yu, Myeongho Jeong, Juneyoung Park, Jamin Shin, Minsam Kim, Seungtaek Choi |
EMNLP | 9 |
| 2021 | Counterfactual Generative Smoothing for Imbalanced Natural Language ClassificationabstractClassification datasets are often biased in observations, leaving onlya few observations for minority classes. Our key contribution is de-tecting and reducing Under-represented (U-) and Over-represented(O-) artifacts from dataset imbalance, by proposing a Counterfac-tual Generative Smoothing approach on both feature-space anddata-space, namely CGS_f and CGS_d. Our technical contribution issmoothing majority and minority observations, by sampling a ma-jority seed and transferring to minority. Our proposed approachesnot only outperform state-of-the-arts in both synthetic and real-lifedatasets, they effectively reduce both artifact types. Hojae Han, Seungtaek Choi, Myeongho Jeong, Seung-won Hwang |
CIKM | 2 |
| 2021 | Structure-Augmented Keyphrase GenerationabstractThis paper studies the keyphrase generation (KG) task for scenarios where structure plays an important role.For example, a scientific publication consists of a short title and a long body, where the title can be used for de-emphasizing unimportant details in the body.Similarly, for short social media posts (e.g., tweets), scarce context can be augmented from titles, though often missing.Our contribution is generating/augmenting structure then encoding these information, using existing keyphrases of other documents, complementing missing/incomplete titles.Specifically, we first extend the given document with related but absent keyphrases from existing keyphrases, to augment missing contexts (generating structure), and then, build a graph of keyphrases and the given document, to obtain structure-aware representation of the augmented text (encoding structure).Our empirical results validate that our proposed structure augmentation and structure-aware encoding can improve KG for both scenarios, outperforming the state-of-the-art 1 . Jihyuk Kim, Myeongho Jeong, Seungtaek Choi, Seung-won Hwang |
EMNLP (1) | 3 |
| 2021 | Label and Context Augmentation for Response Selection at DSTC8abstractThis paper studies the dialogue response selection task. As state-of-the-arts are neural models requiring a large training set, data augmentation has been considered as a means to overcome the sparsity of observational annotation, where only one observed response is annotated as gold. In this paper, we first consider label augmentation, of selecting, among unobserved utterances, that would “counterfactually” replace the labeled response, for the given context, and augmenting labels only if that is the case. The key advantage of this model is not incurring human annotation overhead, thus not increasing the training cost, i.e., for low-resource scenarios. In addition, we consider context augmentation scenarios where the given dialogue context is not sufficient for label augmentation. In this case, inspired by open-domain question answering, we “decontextualize” by retrieving missing contexts, such as related persona. We empirically show that our pipeline improves BERT-based models in two different response selection tasks without incurring annotation overheads. Myeongho Jeong, Seungtaek Choi, Jinyoung Yeo, Seung-won Hwang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Retrieval-Augmented Controllable Review GenerationabstractIn this paper, we study review generation given a set of attribute identifiers which are user ID, product ID and rating.This is a difficult subtask of natural language generation since models are limited to the given identifiers, without any specific descriptive information regarding the inputs, when generating the text.The capacity of these models is thus confined and dependent to how well the models can capture vector representations of attributes.We thus propose to additionally leverage references, which are selected from a large pool of texts labeled with one of the attributes, as textual information that enriches inductive biases of given attributes.With these references, we can now pose the problem as an instance of text-to-text generation, which makes the task easier since texts that are syntactically, semantically similar to the output text are provided as inputs.Using this framework, we address issues such as selecting references from a large candidate set without textual context and improving the model complexity for generation.Our experiments show that our models improve over previous approaches on both automatic and human evaluation metrics. Jihyeok Kim, Seungtaek Choi, Reinald Kim Amplayo, Seung-won Hwang |
COLING | 2 |
| 2020 | Less is More: Attention Supervision with Counterfactuals for Text ClassificationabstractWe aim to leverage human and machine intelligence together for attention supervision.Specifically, we show that human annotation cost can be kept reasonably low, while its quality can be enhanced by machine selfsupervision.Specifically, for this goal, we explore the advantage of counterfactual reasoning, over associative reasoning typically used in attention supervision.Our empirical results show that this machine-augmented human attention supervision is more effective than existing methods requiring a higher annotation cost, in text classification tasks, including sentiment analysis and news categorization. Seungtaek Choi, Haeju Park, Jinyoung Yeo, Seung-won Hwang |
EMNLP (1) | 1 |
| 2020 | Conditional Response Augmentation for Dialogue Using Knowledge Distillation
Myeongho Jeong, Seungtaek Choi, Hojae Han, Kyungho Kim, Seung-won Hwang |
INTERSPEECH | 2 |
| 2020 | Meta-supervision for Attention Using Counterfactual EstimationabstractAbstract Neural attention mechanism has been used as a form of explanation for model behavior. Users can either passively consume explanation or actively disagree with explanation and then supervise attention into more proper values (attention supervision). Though attention supervision was shown to be effective in some tasks, we find the existing attention supervision is biased, for which we propose to augment counterfactual observations to debias and contribute to accuracy gains. To this end, we propose a counterfactual method to estimate such missing observations and debias the existing supervisions. We validate the effectiveness of our counterfactual supervision on widely adopted image benchmark datasets: CUFED and PEC. Seungtaek Choi, Haeju Park, Seung-won Hwang |
Data Sci. Eng. | 1 |
| 2019 | MICRON: Multigranular Interaction for Contextualizing RepresentatiON in Non-factoid Question AnsweringabstractHojae Han, Seungtaek Choi, Haeju Park, Seung-won Hwang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Hojae Han, Seungtaek Choi, Haeju Park, Seung-won Hwang |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Counterfactual Attention SupervisionabstractNeural attention mechanism has been used as a form of explanation for model behavior. Users can either passively consume explanation, or actively disagree with explanation then supervise attention into more proper values (attention supervision). Though attention supervision was shown to be effective in some tasks, we find the existing attention supervision is biased, for which we propose to augment counterfactual observations to debias and contribute to accuracy gains. To this end, we propose a counterfactual method to estimate such missing observations and debias the existing supervisions. We validate the effectiveness of our counterfactual supervision on widely adopted image benchmark datasets: CUFED and PEC. Seungtaek Choi, Haeju Park, Seung-won Hwang |
ICDM | 1 |
| 2018 | Machine-Translated Knowledge Transfer for Commonsense Causal ReasoningabstractThis paper studies the problem of multilingual causal reasoning in resource-poor languages. Existing approaches, translating into the most probable resource-rich language such as English, suffer in the presence of translation and language gaps between different cultural area, which leads to the loss of causality. To overcome these challenges, our goal is thus to identify key techniques to construct a new causality network of cause-effect terms, targeted for the machine-translated English, but without any language-specific knowledge of resource-poor languages. In our evaluations with three languages, Korean, Chinese, and French, our proposed method consistently outperforms all baselines, achieving up-to 69.0% reasoning accuracy, which is close to the state-of-the-art accuracy 70.2% achieved on English. Jinyoung Yeo, Hyunsouk Cho, Seungtaek Choi, Seung-won Hwang |
AAAI | 4 |
| 2018 | Visual Choice of Plausible Alternatives: An Evaluation of Image-based Commonsense Causal Reasoning
Jinyoung Yeo, Gyeongbok Lee, Seungtaek Choi, Hyunsouk Cho, Reinald Kim Amplayo, Seung-won Hwang |
LREC | 4 |