VLDB 2026 Research / reviewers in the wild / expert
Kyohoon Jin
dblp:260/1572
· DBLP profile ↗
13ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0002-7824-3577ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MHTS: Multi-Hop Tree Structure Framework for Generating Difficulty-Controllable QA Datasets for RAG EvaluationabstractExisting RAG benchmarks often overlook query difficulty, leading to inflated performance on simpler questions and unreliable evaluations. A robust benchmark dataset must satisfy three key criteria: quality, diversity, and difficulty, which capturing the complexity of reasoning based on hops and the distribution of supporting evidence. In this paper, we propose MHTS (Multi-Hop Tree Structure), a novel dataset synthesis framework that systematically controls multi-hop reasoning complexity by leveraging a multi-hop tree structure to generate logically connected, multi-chunk queries. Our fine-grained difficulty estimation formula exhibits a strong correlation with the overall performance metrics of a RAG system, validating its effectiveness in assessing both retrieval and answer generation capabilities. By ensuring high-quality, diverse, and difficulty-controlled queries, our approach enhances RAG evaluation and benchmarking capabilities. Jeongsoo Lee, Daeyong Kwon, Kyohoon Jin, Junnyeong Jeong, Minwoo Sim |
LREC | 3 |
| 2025 | SummPilot: Bridging Efficiency and Customization for Interactive Summarization SystemabstractThis paper incorporates the efficiency of automatic summarization and addresses the challenge of generating personalized summaries tailored to individual users' interests and requirements. To tackle this challenge, we introduce SummPilot, an interaction-based customizable summarization system. SummPilot leverages a large language model to facilitate both automatic and interactive summarization. Users can engage with the system to understand document content and personalize summaries through interactive components such as semantic graphs, entity clustering, and explainable evaluation. Our demo and user studies demonstrate SummPilot's adaptability and usefulness for customizable summarization. Jungmin Yun, Juhwan Choi, Kyohoon Jin, Soojin Jang, Jinhee Jang |
AAAI | 3 |
| 2025 | Plug-in and Fine-tuning: Bridging the Gap between Small Language Models and Large Language ModelsabstractLarge language models (LLMs) are renowned for their extensive linguistic knowledge and strong generalization capabilities, but their high computational demands make them unsuitable for resource-constrained environments. In contrast, small language models (SLMs) are computationally efficient but often lack the broad generalization capacity of LLMs. To bridge this gap, we propose PiFi, a novel framework that combines the strengths of both LLMs and SLMs to achieve high performance while maintaining efficiency. PiFi integrates a single frozen layer from an LLM into a SLM and fine-tunes the combined model for specific tasks, boosting performance without a significant increase in computational cost. We show that PiFi delivers consistent performance improvements across a range of natural language processing tasks, including both natural language understanding and generation. Moreover, our findings demonstrate PiFi’s ability to effectively leverage LLM knowledge, enhancing generalization to unseen domains and facilitating the transfer of linguistic abilities. Kyeonghyun Kim, Jinhee Jang, Juhwan Choi, Kyohoon Jin |
ACL (1) | 5 |
| 2025 | CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic TriplesabstractDeep learning models often learn and exploit spurious correlations in training data, using these non-target features to inform their predictions.Such reliance leads to performance degradation and poor generalization on unseen data.To address these limitations, we introduce a more general form of counterfactual data augmentation, termed counterbias data augmentation, which simultaneously tackles multiple biases (e.g., gender bias, simplicity bias) and enhances out-of-distribution robustness.We present COBA: CounterBias Augmentation, a unified framework that operates at the semantic triple level: first decomposing text into subjectpredicate-object triples, then selectively modifying these triples to disrupt spurious correlations.By reconstructing the text from these adjusted triples, COBA generates counterbias data that mitigates spurious patterns.Through extensive experiments, we demonstrate that COBA not only improves downstream task performance, but also effectively reduces biases and strengthens out-of-distribution resilience, offering a versatile and robust solution to the challenges posed by spurious correlations. Kyohoon Jin, Juhwan Choi, Jungmin Yun, Soojin Jang |
EMNLP | 1 |
| 2025 | Korean football in-game conversation state tracking dataset for dialogue and turn level evaluation
Sangmin Song, Juhyoung Park, Juhwan Choi, Kyohoon Jin |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | Enhancing Effectiveness and Robustness in a Low-Resource Regime via Decision-Boundary-aware Data AugmentationabstractEfforts to leverage deep learning models in low-resource regimes have led to numerous augmentation studies. However, the direct application of methods, such as mixup and cutout, is limited due to the discrete characteristics of the textual data. While methods using pre trained language models have exhibited good efficiency, they require additional considerations for robustness. Inspired by recent studies on decision boundaries, this paper proposes a decision-boundary-aware data augmentation strategy to enhance robustness using pretrained language models. The proposed technique first focuses on shifting the latent features closer to the decision boundary, followed by reconstruction to generate an ambiguous version with a soft label. Additionally, mid-K sampling is suggested to enhance the diversity of the generated sentences. This paper demonstrates the performance of the proposed augmentation strategy compared to other methods through extensive experiments. Furthermore, the ablation study demonstrates the effect of soft labels and mid-K sampling and the extensibility of the method with curriculum data augmentation. Kyohoon Jin, Juhwan Choi, Sangmin Song |
LREC/COLING | 1 |
| 2024 | Multi-News+: Cost-efficient Dataset Cleansing via LLM-based Data AnnotationabstractThe quality of the dataset is crucial for ensuring optimal performance and reliability of downstream task models.However, datasets often contain noisy data inadvertently included during the construction process.Numerous attempts have been made to correct this issue through human annotators.However, hiring and managing human annotators is expensive and time-consuming.As an alternative, recent studies are exploring the use of large language models (LLMs) for data annotation.In this study, we present a case study that extends the application of LLM-based data annotation to enhance the quality of existing datasets through a cleansing strategy.Specifically, we leverage approaches such as chain-of-thought and majority voting to imitate human annotation and classify unrelated documents from the Multi-News dataset, which is widely used for the multi-document summarization task.Through our proposed cleansing method, we introduce an enhanced MULTI-NEWS + .By employing LLMs for data cleansing, we demonstrate an efficient and effective approach to improving dataset quality without relying on expensive human annotation efforts. Juhwan Choi, Jungmin Yun, Kyohoon Jin |
EMNLP | 3 |
| 2024 | See, caption, cluster: Large-scale image analysis using captioning and topic modeling
Kyeongpil Kang, Kyohoon Jin, Soojin Jang, Jaegul Choo |
Expert Syst. Appl. | 2 |
| 2023 | Weakly supervised semantic segmentation via Graph RecalibratiOn with Scaling Weight uNit
Soojin Jang, Junehyoung Kwon, Kyohoon Jin |
Eng. Appl. Artif. Intell. | 3 |
| 2023 | Development of robust detector using the weather deep generative model for outdoor monitoring system
Kyohoon Jin, Kyung-Su Kang, Baek-Kyun Shin, Junehyoung Kwon, Soojin Jang, Han Guk Ryu |
Expert Syst. Appl. | 1 |
| 2023 | Game effect sprite generation with minimal data via conditional GAN
Kyohoon Jin, Soojin Jang |
Expert Syst. Appl. | 2 |
| 2021 | Restoring and Mining the Records of the Joseon Dynasty via Neural Language Modeling and Machine TranslationabstractKyeongpil Kang, Kyohoon Jin, Soyoung Yang, Soojin Jang, Jaegul Choo, Youngbin Kim. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Kyeongpil Kang, Kyohoon Jin, Soyoung Yang, Soojin Jang, Jaegul Choo |
NAACL-HLT | 2 |
| 2021 | TrafficBERT: Pre-trained model with large-scale data for long-range traffic flow forecasting
Kyohoon Jin, Jeong A. Wi, Eunju Lee 0003, Sookyun Kim |
Expert Syst. Appl. | 1 |