Minhyuk Kim

dblp:369/4011 · also Min Hyuk Kim, Min-Hyuk Kim, MinHyuk Kim · DBLP profile ↗
← Back
12ranked-venue papers
8as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 7 first-author · 10 since 2021Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 CLEAR: Cross-Lingual Enhancement in Retrieval via Reverse-training
abstract
Seungyoon Lee, Minhyuk Kim, Seongtae Hong, Youngjoon Jang, Dongsuk Oh, Heuiseok Lim. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Seungyoon Lee, Minhyuk Kim, Seongtae Hong, Youngjoon Jang 0002, Dongsuk Oh, Heuiseok Lim
ACL (1)2
2026 FedPure: Data poisoning attack detection and purification for federated skeleton-based action recognition
Minhyuk Kim, Eungi Lee, Seok Bong Yoo
Inf. Sci.1
2025 X-FLoRA: Cross-modal Federated Learning with Modality-expert LoRA for Medical VQA
abstract
Medical visual question answering (VQA) and federated learning (FL) have emerged as vital approaches for enabling privacy-preserving, collaborative learning across clinical institutions. However, both these approaches face significant challenges in cross-modal FL scenarios, where each client possesses unpaired images from only one modality. To address this limitation, we propose X-FLoRA, a cross-modal FL framework that uses modality-expert low-rank adaptation (LoRA) for medical VQA. Specifically, X-FLoRA enables the synthesis of images from one modality to another without requiring data sharing between clients. This is achieved by training a backward translation model within a federated asymmetric translation scheme that integrates clinical semantics from textual data. Additionally, X-FLoRA introduces modality-expert LoRA, which fine-tunes separate LoRA modules to strengthen modality-specific representations in the VQA task. The server aggregates the trained backward translation models and fine-tuned LoRA modules using discriminator quality scores and expert-aware weighting, which regulate the relative contributions from different clients. Experiments were conducted on VQA datasets encompassing different medical modalities, and the results demonstrate that X-FLoRA outperforms existing FL methods in terms of VQA performance.
Minhyuk Kim, Changheon Kim, Seok Bong Yoo
EMNLP1
2025 TORSO: Template-Oriented Reasoning Towards General Tasks
abstract
The approaches that guide Large Language Models (LLMs) to emulate human reasoning during response generation have emerged as an effective method for enabling them to solve complex problems in a step-by-step manner, thereby achieving superior performance.However, most existing approaches using few-shot prompts to generate responses heavily depend on the provided examples, limiting the utilization of the model's inherent reasoning capabilities.Moreover, constructing task-specific few-shot prompts is often costly and may lead to inconsistencies across different tasks.In this work, we introduce Template-Oriented Reasoning (TORSO), which elicits the model to utilize internal reasoning abilities to generate proper responses across various tasks without the need for manually crafted few-shot examples.Our experimental results demonstrate that TORSO achieves strong performance on diverse LLMs benchmarks with reasonable rationales.
Minhyuk Kim, Seungyoon Lee, Heuiseok Lim
EMNLP1
2025 Benchmark Profiling: Mechanistic Diagnosis of LLM Benchmarks
abstract
Large Language Models are commonly judged by their scores on standard benchmarks, yet such scores often overstate real capability since they mask the mix of skills a task actually demands.For example, ARC is assumed to test reasoning, while HellaSwag is designed to evaluate commonsense.However, we lack a systematic way to verify if these benchmarks actually measure these labels.We introduce BENCHMARK PROFILING, a diagnostic framework that decomposes benchmark performance into ten cognitively grounded abilities.The method combines gradient-based importance scoring with targeted parameter ablation to compute an Ability Impact Score (AIS) that quantifies how much each ability contributes to a model's success on a given benchmark.Profiling three instruction-tuned models across ten widely used benchmarks yields four key findings: (i) most benchmarks draw on several abilities rather than one, (ii) datasets with similar labels rely on distinct ability mixtures, (iii) code-generation benchmarks reward broad, multi-skill improvement and thus show only modest gains from narrow domain-specific fine-tuning, and (iv) abilities irrelevant to the task could negatively affect performance.BENCHMARK PROFILING therefore explains why performance gains do not always translate into user-perceived competence and offers a transparent tool for benchmark audit and model interpretability.The code is available on https://github.com/ junkim100/Benchmark-Profiling Leonard Bereska and Efstratios Gavves.
Gyuho Shim, Yongchan Chun, Minhyuk Kim, Chanjun Park, Heuiseok Lim
EMNLP4
2025 Ego-$A^{\mathbf{3}}$: Adaptive Fusion-Based Disentangled Transformer for Egocentric Action Anticipation
abstract
Recently, egocentric action anticipation for wearable robotics cameras has gained considerable attention due to its capability to analyze nouns and verbs from a firstperson view. However, this field encounters challenges due to various uncertainties, such as action-irrelevant information and semantically fused representations of verbs and nouns. To overcome these issues, we introduce Ego-$A^{3}$, designed to improve the robustness and reliability of egocentric action anticipation systems. Ego-$A^{3}$adaptively extracts actionrelevant data to efficiently utilize additional information beyond visual data. Additionally, Ego-$A^{3}$produces effective disentangled representations for verbs and nouns by employing learnable verb and noun queries. Experiments on the EpicKitchens-100 and EGTEA Gaze+ datasets demonstrate that Ego-$A^{3}$outperforms existing methods in top-1 accuracy and mean top- 5 recall. Our code is publicly available at https://github.com/alsgur0720/egocentricanticipation.
Minhyuk Kim, Jong Won Jung, Eungi Lee, Seok Bong Yoo
ICRA1
2025 FedDet: Data Poisoning Attack Detection for Federated Skeleton-based Action Recognition
Minhyuk Kim, Eungi Lee, Seok Bong Yoo
ICRA1
2025 Data Poisoning Attack Defense and Evolutionary Domain Adaptation for Federated Medical Image Segmentation
abstract
Federated learning has significant demonstrated potential in medical image segmentation to protect data privacy by retaining local data. However, its application is still hindered by two critical challenges: 1) the retained data poisoning attacks that severely compromise the accuracy of the global segmentation model and 2) domain gaps among clients, undermining its generalizability. To address these issues, we propose AdaShield-FL, a data poisoning attack defense and evolutionary domain adaptation for federated medical image segmentation. AdaShield-FL incorporates a disentangled reconstruction and segmentation module that purifies data in the k-space domain to mitigate the effects of adversarial attacks iteratively. Moreover, it introduces a data poisoning attack detection mechanism that analyzes abnormal patterns in training loss sequences to identify malicious clients. This method also aligns local and global covariance matrices via evolutionary optimization to minimize the domain gap efficiently. The experimental validation on cardiac magnetic resonance imaging datasets demonstrates the robustness and superior performance of AdaShield-FL compared with other federated learning methods.
Minhyuk Kim, Seok Bong Yoo
IJCAI1
2025 Disentangled adaptive fusion transformer using adversarial perturbation for egocentric action anticipation
Minhyuk Kim, Jong Won Jung, Eungi Lee, Seok Bong Yoo
Expert Syst. Appl.1
2024 Occluded Part-aware Graph Convolutional Networks for Skeleton-based Action Recognition
abstract
Recognizing human action is one of the most critical factors in the visual perception of robots. Specifically, skeletonbased action recognition has been actively researched to enhance recognition performance at a lower cost. However, action recognition in occlusion situations, where body parts are not visible, is still challenging.We propose an occluded part-aware graph convolutional network (OP-GCN) to address this challenge using the optimal occluded body parts. The proposed model uses an occluded part detector to identify occluded body parts within a human skeleton. It is based on an autoencoder trained on a nonoccluded human skeleton and exploits the symmetry and angular information of the skeleton. Then, we select an optimal group constructed considering the occluded body parts. Each group comprises five sets of joint nodes, focusing on the body parts, excluding the occluded ones. Finally, to enhance interaction within the selected groups, we apply an interpart association module, considering the fusion of global and local elements. The experimental results reveal that the proposed model outperforms others on the occluded datasets. These comparative experiments demonstrate the effectiveness of the study in addressing the challenge of action recognition in occlusion situations. Our code is publicly available at https://github.com/MJ-Kor/OP-GCN.
Minhyuk Kim, Seok Bong Yoo
ICRA1
2024 Kernel adaptive memory network for blind video super-resolution
Jun-Seok Yun, Minhyuk Kim, Hyungil Kim, Seok Bong Yoo
Expert Syst. Appl.2
2011 Rate-compatible turbo product codes with non-symmetry block codes for DVB-RCS NG systems
abstract
Soft based linear block codes combined with linear modulation cannot obtain the coding gains by increasing iterations. In this paper, we propose rate-compatible turbo product codes with extended BCH codes by zero padding and shortening the row and/or column to adapt next generation (NG) DVB-RCS system. And also, in order to make various coding rates required on standard, we used different coding rates between row and column codes. At the bit error rate (BER) of 10−4the performance is better around 1.5 [dB] ∼2 [dB] compared to soft based extended BCH codes at the same coding rates.
Tae-Doo Park, Byeong Su Lim, Minhyuk Kim, Ji-Won Jung
APCC3