EDBT 2026 Demo / reviewers in the wild / expert
Jaein Kim 0004
dblp:27/9295-4
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0001-7148-4346ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PeriUn: Enhancing Unlearning by Selectively Forgetting Peripheral SamplesabstractOnce trained, neural networks memorize information in diffusely encoded parameters, making it difficult to forget in support of the right to be forgotten. Unlearning aims to remove the influence of data, with performance measured against a retrained model that excludes the data. However, understanding the behavior of gold-standard retraining remains underexplored. We compare original and retrained models and observe that most prediction changes occur in peripheral samples near decision boundaries. Consequently, we propose PeriUn, a selective strategy that unlearns only peripheral samples to mimic retrained model behavior with minimal disruption, unlike prior works that remove the entire request. Combined with the Random Label based method, PeriUn significantly improves both generalization and privacy metrics. Specifically, on TinyImageNet with VGG16, PeriUn increases the Tug-of-War score by 22 points compared to the strongest. Besides, the MIA gap score surpasses the state-of-the-art method, improving by 8.7 points after applying PeriUn. Further analyses confirm that PeriUn better preserves the feature space and aligns closely with the retrained model. Hee Bin Yoo, Dong-Sig Han, Jaein Kim 0004, Byoung-Tak Zhang |
AAAI | 3 |
| 2025 | How Classifier Features Transfer to Downstream: An Asymptotic Analysis in a Two-Layer ModelabstractNeural networks learn effective feature representations, which can be transferred to new tasks without additional training.
While larger datasets are known to improve feature transfer, the theoretical conditions for the success of such transfer remain unclear.
This work investigates feature transfer in networks trained for classification to identify the conditions that enable effective clustering in unseen classes.
We first reveal that higher similarity between training and unseen distributions leads to improved Cohesion and Separability.
We then show that feature expressiveness is enhanced when inputs are similar to the training classes, while the features of irrelevant inputs remain indistinguishable.
We validate our analysis on synthetic and benchmark datasets, including CAR, CUB, SOP, ISC, and ImageNet.
Our analysis highlights the importance of the similarity between training classes and the input distribution for successful feature transfer. Hee Bin Yoo, Sungyoon Lee, Cheongjae Jang, Dong-Sig Han, Jaein Kim 0004, Seunghyeon Lim, Byoung-Tak Zhang |
NeurIPS | 5 |
| 2024 | Continuous SO(3) Equivariant Convolution for 3D Point Cloud Analysis
Jaein Kim 0004, Hee Bin Yoo, Dong-Sig Han, Yeon-Ji Song, Byoung-Tak Zhang |
ECCV (52) | 1 |
| 2024 | PROGrasp: Pragmatic Human-Robot Communication for Object GraspingabstractInteractive Object Grasping (IOG) is the task of identifying and grasping the desired object via human-robot natural language interaction. Current IOG systems assume that a human user initially specifies the target object’s category (e.g., bottle). Inspired by pragmatics, where humans often convey their intentions by relying on context to achieve goals, we introduce a new IOG task, Pragmatic-IOG, and the corresponding dataset, Intention-oriented Multi-modal Dialogue (IM-Dial). In our proposed task scenario, an intention-oriented utterance (e.g., "I am thirsty") is initially given to the robot. The robot should then identify the target object by interacting with a human user. Based on the task setup, we propose a new robotic system that can interpret the user’s intention and pick up the target object, Pragmatic Object Grasping (PROGrasp). PROGrasp performs Pragmatic-IOG by incorporating modules for visual grounding, question asking, object grasping, and most importantly, answer interpretation for pragmatic inference. Experimental results show that PROGrasp is effective in offline (i.e., target object discovery) and online (i.e., IOG with a physical robot arm) settings. Code and data are available at https://github.com/gicheonkang/prograsp. Gi-Cheon Kang, Junghyun Kim 0009, Jaein Kim 0004, Byoung-Tak Zhang |
ICRA | 3 |
| 2024 | PGA: Personalizing Grasping Agents with Single Human-Robot InteractionabstractLanguage-Conditioned Robotic Grasping (LCRG) aims to develop robots that comprehend and grasp objects based on natural language instructions. While the ability to understand personal objects like my wallet facilitates more natural interaction with human users, current LCRG systems only allow generic language instructions, e.g., the black-colored wallet next to the laptop. To this end, we introduce a task scenario GraspMine alongside a novel dataset aimed at pinpointing and grasping personal objects given personal indicators via learning from a single human-robot interaction, rather than a large labeled dataset. Our proposed method, Personalized Grasping Agent (PGA), addresses GraspMine by leveraging the unlabeled image data of the user’s environment, called Reminiscence. Specifically, PGA acquires personal object information by a user presenting a personal object with its associated indicator, followed by PGA inspecting the object by rotating it. Based on the acquired information, PGA pseudo-labels objects in the Reminiscence by our proposed label propagation algorithm. Harnessing the information acquired from the interactions and the pseudo-labeled objects in the Reminiscence, PGA adapts the object grounding model to grasp personal objects. This results in significant efficiency while previous LCRG systems rely on resource-intensive human annotations—necessitating hundreds of labeled data to learn my wallet. Moreover, PGA outperforms baseline methods across all metrics and even shows comparable performance compared to the fully-supervised method, which learns from 9k annotated data samples. We further validate PGA’s real-world applicability by employing a physical robot to execute GrsapMine. Code and data are publicly available at https://github.com/JHKim-snu/PGA. Junghyun Kim 0009, Gi-Cheon Kang, Jaein Kim 0004, Seoyun Yang, Minjoon Jung, Byoung-Tak Zhang |
IROS | 3 |
| 2023 | Robust Map Fusion with Visual Attention Utilizing Multi-agent RendezvousabstractThe map fusion for multi-robot simultaneous localization and mapping (SLAM) consistently combines robot maps built independently into the global map. An established approach to map fusion is utilizing rendezvous, which refers to an encounter between multiple agents, to calculate the transformation into the global map. However, previous works using rendezvous have a limitation in that they are unreliable for certain circumstances, where the amount of agent observations or overlapping landmarks is limited. This work proposes a novel map fusion system which robustly fuses local maps in challenging rendezvous that lack shared information. Our system utilizes the single visual perception from rendezvous and estimates the relative pose between agents with the DOPE. Then our scheme transforms local maps with an estimated relative pose and predicts the misalignment from approximated maps by utilizing the attention mechanism of the vision transformer. Comparisons with the Hough transform-based method show that ours is significantly better when the overlap between local maps is insufficient. We also verify the robustness of our system against a similar real-world scenario. Jaein Kim 0004, Dong-Sig Han, Byoung-Tak Zhang |
ICRA | 1 |
| 2023 | GVCCI: Lifelong Learning of Visual Grounding for Language-Guided Robotic ManipulationabstractLanguage-Guided Robotic Manipulation (LGRM) is a challenging task as it requires a robot to understand human instructions to manipulate everyday objects. Recent approaches in LGRM rely on pre-trained Visual Grounding (VG) models to detect objects without adapting to manipulation environments. This results in a performance drop due to a substantial domain gap between the pre-training and real-world data. A straight-forward solution is to collect additional training data, but the cost of human-annotation is extortionate. In this paper, we propose Grounding Vision to Ceaselessly Created Instructions (GVCCI), a lifelong learning framework for LGRM, which continuously learns VG without human supervision. GVCCI iteratively generates synthetic instruction via object detection and trains the VG model with the generated data. We validate our framework in offline and online settings across diverse environments on different VG models. Experimental results show that accumulating synthetic data from GVCCI leads to a steady improvement in VG by up to 56.7% and improves resultant LGRM by up to 29.4%. Furthermore, the qualitative analysis shows that the unadapted VG model often fails to find correct objects due to a strong bias learned from the pre-training data. Finally, we introduce a novel VG dataset for LGRM, consisting of nearly 252k triplets of image-object-instruction from diverse manipulation environments. Junghyun Kim 0009, Gi-Cheon Kang, Jaein Kim 0004, Suyeon Shin, Byoung-Tak Zhang |
IROS | 3 |