VLDB 2026 Research / reviewers in the wild / expert
Chanjun Park
dblp:268/1379
· DBLP profile ↗
25ranked-venue papers
4as first author
25since 2021 · last 2026
0000-0002-7200-9632ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 3 first-author · 23 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LangSAE Editing: Improving Multilingual Information Retrieval via Post-hoc Language Identity RemovalabstractDense retrieval in multilingual settings often searches over mixed-language collections, yet multilingual embeddings encode language identity alongside semantics.This language signal can inflate similarity for same-language pairs and crowd out relevant evidence written in other languages.We propose LANGSAE EDIT-ING, a post-hoc sparse autoencoder trained on pooled embeddings that enables controllable removal of language-identity signal directly in vector space.The method identifies languageassociated latent units using cross-language activation statistics, suppresses these units at inference time, and reconstructs embeddings in the original dimensionality, making it compatible with existing vector databases without retraining the base encoder or re-encoding raw text.Experiments across multiple languages show consistent improvements in ranking quality and cross-language coverage, with especially strong gains for script-distinct languages.The LANGSAE model and training code are publicly available.1 Jeongho Yoon, Chanjun Park, Heuiseok Lim |
ACL (1) | 3 |
| 2026 | CEILLM: Confidence-Aware Dynamic Fusion of Collaborative Embedding-Infused Large Language Models for Conversational RecommendationabstractWe introduce a conversational recommender that unifies collaborative filtering signals with a large language model through a Context-Aware Dynamic Embedding Fusion module. A graph-based collaborative filtering backbone learns latent user and item representations, and lightweight adapters inject these embeddings into the language model during fine-tuning. The fusion module uses a learned gate that conditions on dialogue context, model confidence, and user interaction depth to control the contribution of collaborative filtering and language representations at both representation and scoring stages. This adaptive design steers the model toward language knowledge under cold start and toward collaborative filtering when sufficient interactions are available, which yields robust personalization across user regimes. Training follows a multiobjective scheme that promotes confidence-aligned and history-aware gating while preserving the generative and ranking abilities of the language model. We evaluate using Recall@K, NDCG@K, and MRR@K on widely used conversational recommendation datasets and compare against strong collaborative filtering and language baselines. Results show consistent gains with the largest improvements in cold start settings, and ablation and gate analyses indicate meaningful turn-level variation conditioned on dialogue context, uncertainty, and user warmth. Yeo Chan Yoon, Chanjun Park, Jaechoon Jo |
ACM Trans. Inf. Syst. | 2 |
| 2025 | AGENTiGraph: A Multi-Agent Knowledge Graph Framework for Interactive, Domain-Specific LLM ChatbotsabstractAGENTiGraph is a user-friendly, agent-driven system that enables intuitive interaction and management of domain-specific data through the manipulation of knowledge graphs in natural language. It gives non-technical users a complete, visual solution to incrementally build and refine their knowledge bases, allowing multi-round dialogues and dynamic updates without specialized query languages. The flexible design of AGENTiGraph, including intent classification, task planning, and automatic knowledge integration, ensures seamless reasoning between diverse tasks. Evaluated on a 3,500-query benchmark within an educational scenario, the system outperforms strong zero-shot baselines (achieving 95.12% classification accuracy, 90.45% execution success), indicating potential scalability to compliance-critical or multi-step queries in legal and medical domains, e.g., incorporating new statutes or research on the fly. Our open-source demo offers a powerful new paradigm for multi-turn enterprise knowledge management that bridges LLMs and structured graphs. Xinjie Zhao 0004, Moritz Blum, Yingjian Chen, Boming Yang, Luis Marquez-Carpintero, Monica Pina-Navarro, Yanran Fu, So Morikawa, Yusuke Iwasawa, Yutaka Matsuo, Chanjun Park, Irene Li |
CIKM | 12 |
| 2025 | HealthGenie: A Knowledge-Driven LLM Framework for Tailored Dietary GuidanceabstractSeeking dietary guidance often requires navigating complex nutritional knowledge while considering individual health needs. To address this, we present HealthGenie, an interactive platform that leverages the interpretability of knowledge graphs (KGs) and the conversational power of large language models (LLMs) to deliver tailored dietary recommendations alongside integrated nutritional visualizations for fast, intuitive insights. Upon receiving a user query, HealthGenie performs intent refinement and maps user's needs to a curated nutritional knowledge graph. The system then retrieves and visualizes relevant subgraphs, while offering detailed, explainable recommendations. Users can interactively adjust preferences to further tailor results. A within-subject study and quantitative analysis show that HealthGenie reduces cognitive load and interaction effort while supporting personalized, health-aware decision-making. Xinjie Zhao 0004, Ding Xia, Zhongyi Zhou, Rui Yang 0016, Jinghui Lu, Chanjun Park, Irene Li |
CIKM | 8 |
| 2025 | Representing the Under-Represented: Cultural and Core Capability Benchmarks for Developing Thai Large Language ModelsabstractThe rapid advancement of large language models (LLMs) has highlighted the need for robust evaluation frameworks that assess their core capabilities, such as reasoning, knowledge, and commonsense, leading to the inception of certain widely-used benchmark suites such as the H6 benchmark. However, these benchmark suites are primarily built for the English language, and there exists a lack thereof for under-represented languages, in terms of LLM development, such as Thai. On the other hand, developing LLMs for Thai should also include enhancing the cultural understanding as well as core capabilities. To address these dual challenge in Thai LLM research, we propose two key benchmarks: Thai-H6 and Thai Cultural and Linguistic Intelligence Benchmark (ThaiCLI). Through a thorough evaluation of various LLMs with multi-lingual capabilities, we provide a comprehensive analysis of the proposed benchmarks and how they contribute to Thai LLM development. Furthermore, we have made both the datasets and evaluation code publicly available to encourage further research and development for Thai LLMs. Dahyun Kim 0001, Sukyung Lee, Attapol Rutherford, Chanjun Park |
COLING | 5 |
| 2025 | Mixture-of-Clustered-Experts: Advancing Expert Specialization and Generalization in Instruction TuningabstractA sparse Mixture-of-Experts (MoE) architecture has emerged as a highly scalable solution by conditionally activating sub-modules without a proportional increase in computational costs.However, improving expert specialization to enhance performance and generalization remains a challenge for MoE, especially in instruction tuning scenarios characterized by significant input heterogeneity.In this work, we propose the Mixture-of-Clustered-Experts (MoCE) to address this limitation through a dual-stage routing mechanism.The first stage in the mechanism performs expert group routing based on sequence-level features, while the second stage activates the top-k experts within the group at the token level.This approach enables the effective partitioning of heterogeneous inputs based on their knowledge requirements, encouraging expert group specialization while maintaining the advantages of token-level routing.We evaluate MoCE across a comprehensive set of benchmarks, demonstrating its consistent superiority over strong baselines and its enhanced generalization capabilities.Detailed analysis further highlights the robustness and effectiveness of MoCE. Sugyeong Eo, Jung Jun Lee, Chanjun Park, Heuiseok Lim |
EMNLP | 3 |
| 2025 | Benchmark Profiling: Mechanistic Diagnosis of LLM BenchmarksabstractLarge Language Models are commonly judged by their scores on standard benchmarks, yet such scores often overstate real capability since they mask the mix of skills a task actually demands.For example, ARC is assumed to test reasoning, while HellaSwag is designed to evaluate commonsense.However, we lack a systematic way to verify if these benchmarks actually measure these labels.We introduce BENCHMARK PROFILING, a diagnostic framework that decomposes benchmark performance into ten cognitively grounded abilities.The method combines gradient-based importance scoring with targeted parameter ablation to compute an Ability Impact Score (AIS) that quantifies how much each ability contributes to a model's success on a given benchmark.Profiling three instruction-tuned models across ten widely used benchmarks yields four key findings: (i) most benchmarks draw on several abilities rather than one, (ii) datasets with similar labels rely on distinct ability mixtures, (iii) code-generation benchmarks reward broad, multi-skill improvement and thus show only modest gains from narrow domain-specific fine-tuning, and (iv) abilities irrelevant to the task could negatively affect performance.BENCHMARK PROFILING therefore explains why performance gains do not always translate into user-perceived competence and offers a transparent tool for benchmark audit and model interpretability.The code is available on https://github.com/ junkim100/Benchmark-Profiling Leonard Bereska and Efstratios Gavves. Gyuho Shim, Yongchan Chun, Minhyuk Kim, Chanjun Park, Heuiseok Lim |
EMNLP | 5 |
| 2025 | MultiDocFusion : Hierarchical and Multimodal Chunking Pipeline for Enhanced RAG on Long Industrial DocumentsabstractRAG-based QA has emerged as a powerful method for processing long industrial documents.However, conventional text chunking approaches often neglect the complex structures of long industrial documents, causing information loss and reduced answer quality.To address this, we introduce MultiDocFusion, a multimodal chunking pipeline that integrates: (i) detection of document regions using visionbased document parsing, (ii) text extraction from these regions via OCR, (iii) reconstruction of document structure into a hierarchical tree using large language model (LLM)based document section hierarchical parsing (DSHP-LLM), and (iv) construction of hierarchical chunks through DFS-based Grouping.Extensive experiments across industrial benchmarks demonstrate that MultiDocFusion improves retrieval precision by 8-15% and ANLS QA scores by 2-3% compared to baselines, emphasizing the critical role of explicitly leveraging document hierarchy for multimodal document-based QA.These significant performance gains underscore the necessity of structure-aware chunking in enhancing the fidelity of RAG-based QA systems. Joongmin Shin, Chanjun Park, Jeongbae Park, Jaehyung Seo, Heuiseok Lim |
EMNLP | 2 |
| 2025 | LCIRC: A Recurrent Compression Approach for Efficient Long-form Context and Query Dependent Modeling in LLMsabstractSumin An, Junyoung Sung, Wonpyo Park, Chanjun Park, Paul Hongsuck Seo. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Sumin An, Junyoung Sung, Wonpyo Park, Chanjun Park, Hongsuck Seo |
NAACL (Long Papers) | 4 |
| 2025 | CoME: An Unlearning-based Approach to Conflict-free Model EditingabstractDahyun Jung, Jaehyung Seo, Jaewook Lee, Chanjun Park, Heuiseok Lim. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Dahyun Jung, Jaehyung Seo, Jaewook Lee 0008, Chanjun Park, Heuiseok Lim |
NAACL (Long Papers) | 4 |
| 2025 | An analysis on language transfer of pre-trained language model with cross-lingual post-training
Suhyune Son, Chanjun Park, Jungseob Lee, Midan Shim, Chanhee Lee 0004, Yoonna Jang, Jaehyung Seo, Jungwoo Lim, Heuiseok Lim |
Expert Syst. Appl. | 2 |
| 2024 | Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 BenchmarkabstractChanjun Park, Hyeonwoo Kim, Dahyun Kim, SeongHwan Cho, Sanghoon Kim, Sukyung Lee, Yungi Kim, Hwalsuk Lee. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Chanjun Park, Hyeonwoo Kim, Dahyun Kim 0001, Seonghwan Cho, Sukyung Lee, Hwalsuk Lee |
ACL (1) | 1 |
| 2024 | Detecting Critical Errors Considering Cross-Cultural Factors in English-Korean TranslationabstractRecent machine translation (MT) systems have overcome language barriers for a wide range of users, yet they still carry the risk of critical meaning deviation. Critical error detection (CED) is a task that identifies an inherent risk of catastrophic meaning distortions in the machine translation output. With the importance of reflecting cultural elements in detecting critical errors, we introduce the culture-aware “Politeness” type in detecting English-Korean critical translation errors. Besides, we facilitate two tasks by providing multiclass labels: critical error detection and critical error type classification (CETC). Empirical evaluations reveal that our introduced data augmentation approach using a newly presented perturber significantly outperforms existing baselines in both tasks. Further analysis highlights the significance of multiclass labeling by demonstrating its superior effectiveness compared to binary labels. Sugyeong Eo, Jungwoo Lim, Chanjun Park, Dahyun Jung, Seonmin Koo, Hyeonseok Moon, Jaehyung Seo, Heuiseok Lim |
LREC/COLING | 3 |
| 2024 | Leveraging Pre-existing Resources for Data-Efficient Counter-Narrative Generation in KoreanabstractCounter-narrative generation, i.e., the generation of fact-based responses to hate speech with the aim of correcting discriminatory beliefs, has been demonstrated to be an effective method to combat hate speech. However, its effectiveness is limited by the resource-intensive nature of dataset construction processes and only focuses on the primary language. To alleviate this problem, we propose a Korean Hate Speech Counter Punch (KHSCP), a cost-effective counter-narrative generation method in the Korean language. To this end, we release the first counter-narrative generation dataset in Korean and pose two research questions. Under the questions, we propose an effective augmentation method and investigate the reasonability of a large language model to overcome data scarcity in low-resource environments by leveraging existing resources. In this regard, we conduct several experiments to verify the effectiveness of the proposed method. Our results reveal that applying pre-existing resources can improve the generation performance by a significant margin. Through deep analysis on these experiments, this work proposes the possibility of overcoming the challenges of generating counter-narratives in low-resource environments. Seungyoon Lee, Chanjun Park, Dahyun Jung, Hyeonseok Moon, Jaehyung Seo, Sugyeong Eo, Heuiseok Lim |
LREC/COLING | 2 |
| 2024 | Where am I? Large Language Models Wandering between Semantics and Structures in Long ContextsabstractAs the utilization of Large Language Models (LLMs) becomes more widespread, there is a growing demand for their ability to handle more complex and longer external knowledge across various use cases.Most existing evaluations of the open-ended question answering (ODQA) task, which necessitates the use of external knowledge, focus solely on whether the model provides the correct answer.However, even when LLMs answer correctly, they often fail to provide an obvious source for their responses.Therefore, it is necessary to jointly evaluate and verify the correctness of the answers and the appropriateness of grounded evidence in complex external contexts.To address this issue, we examine the phenomenon of discrepancies in abilities across two distinct tasks-QA and evidence selection-when performed simultaneously, from the perspective of task alignment.To verify LLMs' task alignment, we introduce a verification framework and resources considering both semantic relevancy and structural diversity of the given long context knowledge.Through extensive experiments and detailed analysis, we provide insights into the task misalignment between QA and evidence selection.Our code and resources can be found at https://github.com/seonminkoo/WAI. Seonmin Koo, Youngjoon Jang 0002, Chanjun Park, Heuiseok Lim |
EMNLP | 4 |
| 2023 | KEBAP: Korean Error Explainable Benchmark Dataset for ASR and Post-processingabstractAutomatic Speech Recognition (ASR) systems are instrumental across various applications, with their performance being critically tied to user satisfaction.Conventional evaluation metrics for ASR systems produce a singular aggregate score, which is insufficient for understanding specific system vulnerabilities.Therefore, we aim to address the limitations of the previous ASR evaluation methods by introducing the Korean Error Explainable Benchmark Dataset for ASR and Post-processing (KEBAP).KE-BAP enables comprehensive analysis of ASR systems at both speech-and text levels, thereby facilitating a more balanced assessment encompassing speech recognition accuracy and user readability.KEBAP provides 37 newly defined speech-level resources incorporating diverse noise environments and speaker characteristics categories, also presenting 13 distinct textlevel error types.This paper demonstrates detailed statistical analyses of colloquial noise categories and textual error types.Furthermore, we conduct extensive validation and analysis on commercially deployed ASR systems, providing valuable insights into their performance.As a more fine-grained and real-world-centric evaluation method, KEBAP contributes to identifying and mitigating potential weaknesses in ASR systems.* Equally contributed, ‡ Corresponding author 1 Recognition accuracy is the measure of accurately perceiving phonemes as they are externally expressed, regardless of user input quality (Liao et al., 2022).Conventional (WER, CER) 0.45 KEBAP Error types Explainability Noise Type Description Washer/dryer machine Home appliances Vacuum cleaner Difficulty in recognition due to ambient electrical appliance noise.Motorcycle Siren Individual transportation Honk Difficulty in recognition due to surrounding individual transportation noise.Road side Street Crowd Difficulty in recognition due to the surrounding street noise.Conversation Cafe/restaurant Non-conversation Challenges in perception due to the noise in cafes/restaurants.Traditional market Market/shopping mall Shopping mall Difficulties in perception caused by the noise in markets/shopping malls.Subway platform Inside the subway Inside the train (STR/KTX) Public transportation Inside the bus Difficulty in recognition due to surrounding public transportation noise.Train terminal waiting room Terminal Bus terminal waiting room Challenges in perception due to the noise at terminals.Outdoor construction site Construction site Indoor construction site Difficulties in perception caused by the noise at construction sites.processing process Factory Assembly process Difficulties in perception caused by the noise in factories.Sound of rain Nature ambient Sound of the waves Challenges in perception due to natural ambient noise.Noisy environment Etc.Artificial mechanical sound In cases where external noise is present, although not falling into the aforementioned categories. Seonmin Koo, Chanjun Park, Jaehyung Seo, Sugyeong Eo, Hyeonseok Moon, Heuiseok Lim |
EMNLP | 2 |
| 2023 | CHEF in the Language Kitchen: A Generative Data Augmentation Leveraging Korean Morpheme IngredientsabstractKorean morphological variations present unique opportunities and challenges in natural language processing (NLP), necessitating an advanced understanding of morpheme-based sentence construction.The complexity of morphological variations allows for diverse sentence forms based on the syntactic-semantic integration of functional morphemes (i.e., affixes) to lexical morphemes (i.e., roots).With this in mind, we propose a method -CHEF, replicating the morphological transformations inherent in sentences based on lexical and functional morpheme combinations through generative data augmentation.CHEF operates using a morpheme blender and a label discriminator, thereby enhancing the diversity of Korean sentence forms by capturing the properties of agglutination while maintaining label consistency.We conduct experiments on Korean multiple classification datasets, improving model performance in full-and few-shot settings.Our proposed method boosts performance beyond the preceding data augmentation methods without incurring external data usage.We demonstrate that our approach achieves comparable results yielded by augmentation techniques that use large language models (LLMs). Jaehyung Seo, Hyeonseok Moon, Jaewook Lee 0008, Sugyeong Eo, Chanjun Park, Heuiseok Lim |
EMNLP | 5 |
| 2023 | Informative Evidence-guided Prompt-based Fine-tuning for English-Korean Critical Error DetectionabstractDaHyun Jung, Sugyeong Eo, Chanjun Park, Hyeonseok Moon, Jaehyung Seo, Heuiseok Lim. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Dahyun Jung, Sugyeong Eo, Chanjun Park, Hyeonseok Moon, Jaehyung Seo, Heuiseok Lim |
IJCNLP (1) | 3 |
| 2023 | Doubts on the reliability of parallel corpus filtering
Hyeonseok Moon, Chanjun Park, Seonmin Koo, Jungseob Lee, Jaehyung Seo, Sugyeong Eo, Yoonna Jang, Hyunjoong Kim, Hyoung-gyu Lee, Heuiseok Lim |
Expert Syst. Appl. | 2 |
| 2022 | QUAK: A Synthetic Quality Estimation Dataset for Korean-English Neural Machine TranslationabstractWith the recent advance in neural machine translation demonstrating its importance, research on quality estimation (QE) has been steadily progressing. QE aims to automatically predict the quality of machine translation (MT) output without reference sentences. Despite its high utility in the real world, there remain several limitations concerning manual QE data creation: inevitably incurred non-trivial costs due to the need for translation experts, and issues with data scaling and language expansion. To tackle these limitations, we present QUAK, a Korean-English synthetic QE dataset generated in a fully automatic manner. This consists of three sub-QUAK datasets QUAK-M, QUAK-P, and QUAK-H, produced through three strategies that are relatively free from language constraints. Since each strategy requires no human effort, which facilitates scalability, we scale our data up to 1.58M for QUAK-P, H and 6.58M for QUAK-M. As an experiment, we quantitatively analyze word-level QE results in various ways while performing statistical analysis. Moreover, we show that datasets scaled in an efficient way also contribute to performance improvements by observing meaningful performance gains in QUAK-M, P when adding data up to 1.58M. Sugyeong Eo, Chanjun Park, Hyeonseok Moon, Jaehyung Seo, Gyeongmin Kim, Jungseob Lee, Heuiseok Lim |
COLING | 2 |
| 2022 | Empirical Analysis of Noising Scheme based Synthetic Data Generation for Automatic Post-editingabstractAutomatic post-editing (APE) refers to a research field that aims to automatically correct errors included in the translation sentences derived by the machine translation system. This study has several limitations, considering the data acquisition, because there is no official dataset for most language pairs. Moreover, the amount of data is restricted even for language pairs in which official data has been released, such as WMT. To solve this problem and promote universal APE research regardless of APE data existence, this study proposes a method for automatically generating APE data based on a noising scheme from a parallel corpus. Particularly, we propose a human mimicking errors-based noising scheme that considers a practical correction process at the human level. We propose a precise inspection to attain high performance, and we derived the optimal noising schemes that show substantial effectiveness. Through these, we also demonstrate that depending on the type of noise, the noising scheme-based APE data generation may lead to inferior performance. In addition, we propose a dynamic noise injection strategy that enables the acquisition of a robust error correction capability and demonstrated its effectiveness by comparative analysis. This study enables obtaining a high performance APE model without human-generated data and can promote universal APE research for all language pairs targeting English. Hyeonseok Moon, Chanjun Park, Seolhwa Lee, Jaehyung Seo, Jungseob Lee, Sugyeong Eo, Heuiseok Lim |
LREC | 2 |
| 2022 | FreeTalky: Don't Be Afraid! Conversations Made Easier by a Humanoid Robot using Persona-based DialogueabstractWe propose a deep learning-based foreign language learning platform, named FreeTalky, for people who experience anxiety dealing with foreign languages, by employing a humanoid robot NAO and various deep learning models. A persona-based dialogue system that is embedded in NAO provides an interesting and consistent multi-turn dialogue for users. Also, an grammar error correction system promotes improvement in grammar skills of the users. Thus, our system enables personalized learning based on persona dialogue and facilitates grammar learning of a user using grammar error feedback. Furthermore, we verified whether FreeTalky provides practical help in alleviating xenoglossophobia by replacing the real human in the conversation with a NAO robot, through human evaluation. Chanjun Park, Yoonna Jang, Seolhwa Lee, Heuiseok Lim |
LREC | 1 |
| 2022 | Priming Ancient Korean Neural Machine TranslationabstractIn recent years, there has been an increasing need for the restoration and translation of historical languages. In this study, we attempt to translate historical records in ancient Korean language based on neural machine translation (NMT). Inspired by priming, a cognitive science theory that two different stimuli influence each other, we propose novel priming ancient-Korean NMT (AKNMT) using bilingual subword embedding initialization with structural property awareness in the ancient documents. Finally, we obtain state-of-the-art results in the AKNMT task. To the best of our knowledge, we confirm the possibility of developing a human-centric model that incorporates the concepts of cognitive science and analyzes the result from the perspective of interference and cognitive dissonance theory for the first time. Chanjun Park, Seolhwa Lee, Jaehyung Seo, Hyeonseok Moon, Sugyeong Eo, Heuiseok Lim |
LREC | 1 |
| 2022 | PU-GEN: Enhancing generative commonsense reasoning for language models with human-centered knowledge
Jaehyung Seo, Dongsuk Oh, Sugyeong Eo, Chanjun Park, Kisu Yang, Hyeonseok Moon, Kinam Park, Heuiseok Lim |
Knowl. Based Syst. | 4 |
| 2021 | Neural spelling correction: translating incorrect sentences to correct sentences for multimedia
Chanjun Park, Kuekyeng Kim, YeongWook Yang, Minho Kang, Heuiseok Lim |
Multim. Tools Appl. | 1 |