Kwangsu Kim

dblp:71/6192 · DBLP profile ↗
← Back
15ranked-venue papers
2as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
3D vision · 33% Optimization for machine learning · 18% Efficient and distributed learning · 18%
Computer graphics and multimedia
2 papers
Computational photography and imaging · 46% Multimedia analysis and retrieval · 27% Visualization and visual analytics · 27%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 77% Energy-efficient computing · 23%

Topics — the 22 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › neural rendering
3d gaussian splatting
1.012026
COSMOS: Coherent SuperGaussian Modeling with Spatial Priors for Sparse-View 3D Splatting · AAAI 2026
Computer vision › 3D vision
3d reconstruction
1.012026
COSMOS: Coherent SuperGaussian Modeling with Spatial Priors for Sparse-View 3D Splatting · AAAI 2026
Computer vision › 3D vision › 3d reconstruction › multi-view reconstruction
sparse-view reconstruction
1.012026
COSMOS: Coherent SuperGaussian Modeling with Spatial Priors for Sparse-View 3D Splatting · AAAI 2026
Machine learning › Trustworthy machine learning › interpretability
concept-based explanation
0.912025
Img2Tab: Automatic Class Relevant Concept Discovery from StyleGAN Features for Explainable Image Classification · Int. J. Comput. Vis. 2025
Machine learning › Transfer learning and domain adaptation
domain generalization
0.912025
One-Step Generalization Ratio Guided Optimization for Domain Generalization · ICML 2025
Machine learning › Efficient and distributed learning
federated learning
0.912025
Federated Learning for Feature Generalization with Convex Constraints · ICML 2025
Machine learning › Efficient and distributed learning › federated learning
federated optimization
0.912025
Federated Learning for Feature Generalization with Convex Constraints · ICML 2025
Machine learning › Optimization for machine learning › multi-task optimization
gradient alignment
0.912025
One-Step Generalization Ratio Guided Optimization for Domain Generalization · ICML 2025
Machine learning › Optimization for machine learning
gradient-based optimization
0.912025
One-Step Generalization Ratio Guided Optimization for Domain Generalization · ICML 2025
Machine learning › Trustworthy machine learning
interpretability
0.912025
Img2Tab: Automatic Class Relevant Concept Discovery from StyleGAN Features for Explainable Image Classification · Int. J. Comput. Vis. 2025
Visualization and visual analytics › visual attention
spatio-temporal attention
0.812024
CSTA: CNN-based Spatiotemporal Attention for Video Summarization · CVPR 2024
Multimedia analysis and retrieval
video summarization
0.812024
CSTA: CNN-based Spatiotemporal Attention for Video Summarization · CVPR 2024
Computational photography and imaging
acoustic imaging
0.712023
Acoustic NLOS Imaging with Cross Modal Knowledge Distillation · IJCAI 2023
Computational photography and imaging
non-line-of-sight imaging
0.712023
Acoustic NLOS Imaging with Cross Modal Knowledge Distillation · IJCAI 2023
Energy-efficient computing › power management › memory power management
cache energy reduction
0.412020
Per-Operation Reusability Based Allocation and Migration Policy for Hybrid Cache · IEEE Trans. Computers 2020
Memory systems › cache › cache technology
hybrid cache
0.412020
Per-Operation Reusability Based Allocation and Migration Policy for Hybrid Cache · IEEE Trans. Computers 2020
Memory systems
cache management
0.312018
Replacement Policy Adaptable Miss Curve Estimation for Efficient Cache Partitioning · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Memory systems › cache management
cache partitioning
0.312018
Replacement Policy Adaptable Miss Curve Estimation for Efficient Cache Partitioning · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Memory systems › cache management
cache replacement
0.312018
Replacement Policy Adaptable Miss Curve Estimation for Efficient Cache Partitioning · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Computer vision › 3D vision › local feature descriptor
superpoint
0.312026
COSMOS: Coherent SuperGaussian Modeling with Spatial Priors for Sparse-View 3D Splatting · AAAI 2026
Computer vision › Image recognition and object detection
image classification
0.312025
Img2Tab: Automatic Class Relevant Concept Discovery from StyleGAN Features for Explainable Image Classification · Int. J. Comput. Vis. 2025
Computer vision › Video understanding and tracking
video summarization
0.212024
CSTA: CNN-based Spatiotemporal Attention for Video Summarization · CVPR 2024

Methods — techniques the papers use, named apart from their topics

self-attention · 1.0positional regularization · 1.0wasserstein-1 metric · 0.9tabular classifier · 0.9preconditioning · 0.9one-step generalization ratio · 0.9gradient signal-to-noise ratio analysis · 0.9dropout · 0.9convex constraints · 0.9StyleGAN inversion · 0.9convolutional neural network · 0.8attention mechanism · 0.8deep learning · 0.7cross-modal knowledge distillation · 0.7cost function · 0.4LRU replacement · 0.4replacement policy adaptable estimation · 0.3
YearPublicationVenuePosition
2026 COSMOS: Coherent SuperGaussian Modeling with Spatial Priors for Sparse-View 3D Splatting
abstract
3D Gaussian Splatting (3DGS) has recently emerged as a promising approach for 3D reconstruction, providing explicit, point-based representations and enabling high-quality real-time rendering. However, when trained with sparse input views, 3DGS suffers from overfitting and structural degradation, leading to poor generalization on novel views. This limitation arises from its optimization relying solely on photometric loss without incorporating any 3D structure priors. To address this issue, we propose Coherent supergaussian Modeling with Spatial Priors (COSMOS). Inspired by the concept of superpoints from 3D segmentation, COSMOS introduces 3D structure priors by newly defining supergaussian groupings of Gaussians based on local geometric cues and appearance features. To this end, COSMOS applies inter-group global self-attention across supergaussian groups and sparse local attention among individual Gaussians, enabling the integration of global and local spatial information. These structure-aware features are then used for predicting Gaussian attributes, facilitating more consistent 3D reconstructions. Furthermore, by leveraging supergaussian-based grouping, COSMOS enforces an intra-group positional regularization to maintain structural coherence and suppress floaters, thereby enhancing training stability under sparse-view conditions. Our experiments on Blender and DTU show that COSMOS surpasses state‑of‑the‑art methods in sparse‑view settings without any external depth supervision.
Chaeyoung Jeong, Kwangsu Kim
AAAI2
2026 TM-Adapter: Temporal Merge Adapter for Efficient Global Temporal Modeling
Woo Joo Hahm, Seungwoo Jang, Hyeon Tak Kim, Kwangsu Kim
WACV5
2026 Label-agnostic, Representation-Level Machine Unlearning via feature disentanglement
abstract
Machine Unlearning (MU) aims to selectively remove the influence of specific data from a trained model, driven by increasing demands for data privacy and legal compliance under regulations such as GDPR and CCPA. However, most existing MU methods rely on explicit label supervision, which is often unavailable or costly to obtain at unlearning time, and may also raise privacy concerns, as labels themselves can reveal sensitive information. Moreover, forget data often contains both information shared with retain data and information specific to itself. Removing shared information can lead to significant performance degradation on the retained knowledge. In response, we propose a Label-Agnostic, Representation-Level Machine Unlearning framework that disentangles and selectively removes only the forget data-specific representations without requiring supervision. Our approach first disentangles the representations of forget and retain data via a dedicated disentanglement objective, then re-initializes the forget-specific components while preserving shared representations. This enables targeted unlearning directly at the representation level without compromising the utility of the retained knowledge. Experimental results demonstrate that our method achieves effective unlearning performance by selectively removing forget-specific information while preserving overall model utility, enabled by a supervision-free framework that performs disentanglement-based representation forgetting and ensures strong privacy preservation under label-agnostic constraints. We report state-of-the-art performance on class-unlearning, sub-class unlearning, random sample unlearning and noisy label entry unlearning across multiple datasets and architectures.
Soobin Cha, Jongmin Lim, Jaehun Park, Kwangsu Kim
Knowl. Based Syst.5
2025 One-Step Generalization Ratio Guided Optimization for Domain Generalization
abstract
Domain Generalization (DG) aims to train models that generalize to unseen target domains but often overfit to domain-specific features, known as undesired correlations. Gradient-based DG methods typically guide gradients in a dominant direction but often inadvertently reinforce spurious correlations. Recent work has employed dropout to regularize overconfident parameters, but has not explicitly adjusted gradient alignment or ensured balanced parameter updates. We propose GENIE (Generalization-ENhancing Iterative Equalizer), a novel optimizer that leverages the One-Step Generalization Ratio (OSGR) to quantify each parameter's contribution to loss reduction and assess gradient alignment. By dynamically equalizing OSGR via a preconditioning factor, GENIE prevents a small subset of parameters from dominating optimization, thereby promoting domain-invariant feature learning. Theoretically, GENIE balances convergence contribution and gradient alignment among parameters, achieving higher OSGR while retaining SGD's convergence rate. Empirically, it outperforms existing optimizers and enhances performance when integrated with various DG and single-DG methods.
Sumin Cho, Kwangsu Kim
ICML3
2025 Federated Learning for Feature Generalization with Convex Constraints
abstract
Federated learning (FL) often struggles with generalization due to heterogeneous client data. Local models are prone to overfitting their local data distributions, and even transferable features can be distorted during aggregation. To address these challenges, we propose FedCONST, an approach that adaptively modulates update magnitudes based on the global model’s parameter strength. This prevents over-emphasizing well-learned parameters while reinforcing underdeveloped ones. Specifically, FedCONST employs linear convex constraints to ensure training stability and preserve locally learned generalization capabilities during aggregation. A Gradient Signal-to-Noise Ratio (GSNR) analysis further validates FedCONST's effectiveness in enhancing feature transferability and robustness. As a result, FedCONST effectively aligns local and global objectives, mitigating overfitting and promoting stronger generalization across diverse FL environments, achieving state-of-the-art performance.
Donghee Kim, Sung Kuk Shyn, Kwangsu Kim
ICML4
2025 Img2Tab: Automatic Class Relevant Concept Discovery from StyleGAN Features for Explainable Image Classification
abstract
Traditional tabular classifiers provide explainable decision-making with interpretable features (concepts). However, using their explainability in vision tasks has been limited due to the pixel representation of images. In this paper, we design Img2Tabs that classify images by concepts to harness the explainability of tabular classifiers. Img2Tabs encode image pixels into tabular features by StyleGAN inversion. Since not all of the resulting features are class-relevant or interpretable due to their generative nature, the Img2Tab classifier should automatically discover class-relevant concepts from the StyleGAN features. Thus, we propose a novel algorithm using the Wasserstein-1 metric to quantify class-relevancy and interpretability simultaneously. By this method of concept visualization, we quantitatively investigate whether important features extracted by tabular classifiers are class-relevant concepts. Consequently, we determine the most effective classifier for Img2Tabs in terms of discovering class-relevant concepts automatically from StyleGAN features. In evaluations, we demonstrate concept-based explanations through importance and visualization. Img2Tab achieves top-1 accuracy on par with CNN classifiers and deep feature learning baselines. Additionally, we show that users can interactively debug Img2Tab classifier to prevent erroneous decision-making from data bias without sacrificing accuracy. The source and demo code for Img2Tab are available at https://github.com/songsnim/Img2Tab_pytorch
Youngjae Song, Sung Kuk Shyn, Kwangsu Kim
Int. J. Comput. Vis.3
2025 Meta-learning with gradient norm arbitration for sample-aware few-shot learning
abstract
The ability to rapidly adapt to unseen tasks is a fundamental objective in few-shot learning. Recent advances in optimization-based meta-learning have enhanced adaptability by learning sharable prior knowledge across tasks with just a few gradient descent steps. However, we argue that this shared prior knowledge can exert an imbalanced influence on individual samples within tasks, potentially resulting in a broad loss distribution where samples closely aligned with the prior knowledge exhibit low loss values, while others display high loss values. Furthermore, our experiments show that gradients computed as the average from a broad loss distribution tend to be non-representative and low, leading to poor generalization performance since the contribution of high-loss samples is diminished by low-loss samples. To address this, we propose a novel meta-learning method that arbitrates gradient norms based on sample-aware information during task adaptation. Specifically, we first normalize the gradient vector to reduce the imbalanced influence of prior knowledge on individual samples. Subsequently, the Arbiter, a learnable network, dynamically scales the current gradient norm by analyzing the relationship between original gradient norms and weight norms, which indicates the model’s sensitivity and complexity to each sample. In this way, the proposed method, Meta-learning with Gradient Norm Arbitration (Meta-GNA), improves generalization performance by preserving more representative and higher gradients that adequately reflect high-loss samples, which are distantly aligned with prior knowledge. Experimental results show that Meta-GNA improves performance in few-shot classification, particularly in cross-domain scenarios where the imbalance in prior knowledge across samples is more pronounced.
Jongmin Lim, Soobin Cha, Heesan Kong, Sung Kuk Shyn, Kwangsu Kim
Knowl. Based Syst.5
2024 CSTA: CNN-based Spatiotemporal Attention for Video Summarization
abstract
Video summarization aims to generate a concise repre-sentation of a video, capturing its essential content and key moments while reducing its overall length. Although several methods employ attention mechanisms to handle long-term dependencies, they often fail to capture the visual signif-icance inherent in frames. To address this limitation, we propose a CNN-based SpatioTemporal Attention (CSTA) method that stacks each feature of frames from a single video to form image-like frame representations and applies 2D CNN to these frame features. Our methodology relies on CNN to comprehend the inter and intra-frame relations and to find crucial attributes in videos by exploiting its abil-ity to learn absolute positions within images. In contrast to previous work compromising efficiency by designing additional modules to focus on spatial importance, CSTA re-quires minimal computational overhead as it uses CNN as a sliding window. Extensive experiments on two benchmark datasets (SumMe and TVSum) demonstrate that our pro-posed approach achieves state-of-the-art performance with fewer MACs compared to previous methods. Codes are available at https://github.com/thswodnjs3/CSTA.
Jaewon Son, Jaehun Park, Kwangsu Kim
CVPR3
2023 Acoustic NLOS Imaging with Cross Modal Knowledge Distillation
abstract
Acoustic non-line-of-sight (NLOS) imaging aims to reconstruct hidden scenes by analyzing reflections of acoustic waves. Despite recent developments in the field, existing methods still have limitations such as sensitivity to noise in a physical model and difficulty in reconstructing unseen objects in a deep learning model. To address these limitations, we propose a novel cross-modal knowledge distillation (CMKD) approach for acoustic NLOS imaging. Our method transfers knowledge from a well-trained image network to an audio network, effectively combining the strengths of both modalities. As a result, it is robust to noise and superior in reconstructing unseen objects. Additionally, we evaluate real-world datasets and demonstrate that the proposed method outperforms state-of-the-art methods in acoustic NLOS imaging. The experimental results indicate that CMKD is an effective solution for addressing the limitations of current acoustic NLOS imaging methods. Our code, model, and data are available at https://github.com/shineh96/Acoustic-NLOS-CMKD.
Ui-Hyeon Shin, Seungwoo Jang, Kwangsu Kim
IJCAI3
2022 Generalized Facial Manipulation Detection with Edge Region Feature Extraction
abstract
This paper presents a generalized and robust face manipulation detection method based on the edge region features appearing in images. Most contemporary face synthesis processes include color awkwardness reduction but damage the natural fingerprint in the edge region. In addition, these color correction processes do not proceed in the non-face background region. We also observe that the synthesis process does not consider the natural properties of the image appearing in the time domain. Considering these observations, we propose a facial forensic framework that utilizes pixel-level color features appearing in the edge region of the whole image. Furthermore, our framework includes a 3D-CNN classification model that interprets the extracted color features spatially and temporally. Unlike other existing studies, we conduct authenticity determination by considering all features extracted from multiple frames within one video. Through extensive experiments, including real-world scenarios to evaluate generalized detection ability, we show that our framework outperforms state-of-the-art facial manipulation detection technologies in terms of accuracy and robustness.
Dong-Keon Kim, Kwangsu Kim
WACV2
2020 Per-Operation Reusability Based Allocation and Migration Policy for Hybrid Cache
abstract
Recently, a hybrid cache consisting of SRAM and STT-RAM has attracted much attention as a future memory by complementing each other with different memory characteristics. Prior works focused on developing data allocation and migration techniques considering write-intensity to reduce write energy at STT-RAM. However, these works often neglect the impact of operation-specific reusability of a cache line. In this paper, we propose an energy-efficient per-operation reusability-based allocation and migration policy (ORAM) with a unified LRU replacement policy. First, to select an adequate memory type for allocation, we propose a cost function based on per-operation reusability - gain from an allocated cache line and loss from an evicted cache line for different memory types - which exploits the temporal locality. Besides, we present a migration policy, victim and target cache line selection scheme, to resolve memory type inconsistency between replacement policy and the allocation policy, with further energy reduction. Experiment results show an average energy reduction in the LLC and the main memory by 12.3 and 21.2 percent, and the improvement of latency and execution time by 21.2 and 8.8 percent, respectively, compared with a baseline hybrid cache management. In addition, the Energy-Delay Product (EDP) is improved by 36.9 percent over the baseline.
Minsik Oh, Kwangsu Kim, Duheon Choi, Hyuk-Jun Lee, Eui-Young Chung
IEEE Trans. Computers2
2018 Replacement Policy Adaptable Miss Curve Estimation for Efficient Cache Partitioning
abstract
Cache replacement policies and cache partitioning are well-known cache management techniques which aim to eliminate inter- and intra-application contention caused by co-running applications, respectively. Since replacement policies can change applications' behavior on a shared last-level cache, they have a massive impact on cache partitioning. Furthermore, cache partitioning determines the capacity allocated to each application affecting incorporated replacement policy. However, their interoperability has not been thoroughly explored. Since existing cache partitioning methods are tailored to specific replacement policies to reduce overheads for characterization of applications' behavior, they may lead to suboptimal partitioning results when incorporated with the up-to-date replacement policies. In cache partitioning, miss curve estimation is a key component to relax this restriction which can reflect the dependency between a replacement policy and cache partitioning on partitioning decision. To tackle this issue, we propose a replacement policy adaptable miss curve estimation (RME) which estimates dynamic workload patterns according to any arbitrary replacement policy and to given applications with low overhead. In addition, RME considers asymmetry of miss latency by miss type, thus the impact of miss curve on cache partitioning can be reflected more accurately. The experimental results support the efficiency of RME and show that RME-based cache partitioning cooperated with high-performance replacement policies can minimize both inter- and intra-application interference successfully.
Byunghoon Lee, Kwangsu Kim, Eui-Young Chung
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2008 Distributed Visual Processing for a Home Visual Sensor Network
abstract
We address issues dealing with distributed visual processing for a personal service robot in the Intelligent Home environment. We propose an efficient and reliable framework to organize and coordinate the vision sensor nodes: fixed cameras mounted on walls, and camera(s) on the mobile robot. We also propose key visual functionalities necessary for the robot to perform its activities. They include people detection and identification, action recognition, gesture recognition, and self-localization. We propose solutions to the different vision tasks, and present our implementation within this framework, validated with experimental results.
Kwangsu Kim, Gérard G. Medioni
WACV1
2007 Robust Real-Time Vision for a Personal Service Robot in a Home Visual Sensor Network
abstract
We address issues dealing with visual perception for a personal service robot in the intelligent home environment. We identify key visual functionalities necessary for the robot to perform its activities. They include people detection and identification, gesture recognition, and self-localization. We propose an efficient and reliable framework to organize and coordinate the vision sensor nodes: fixed cameras mounted on walls, and earnera(s) on the mobile robot. We propose solutions to the different vision tasks, and present our implementation within this framework, validated with experimental results.
Kwangsu Kim, Gérard G. Medioni
RO-MAN1
2007 Robust real-time vision for a personal service robot
Gérard G. Medioni, Alexandre R. J. François, Matheen Siddiqui, Kwangsu Kim, Hosub Yoon
Comput. Vis. Image Underst.4